We are looking for a seasoned DevOps SRE Specialist who brings deep, hands-on expertise across both cloud infrastructure engineering and site reliability. This is a hybrid role that requires someone equally comfortable designing Kubernetes-based platforms on AWS and Azure, as they are building observability pipelines and responding to production incidents.
You will own the full lifecycle of platform infrastructure — from cluster provisioning and auto-scaling to distributed tracing, on-call incident management, and continuous delivery pipelines. You will work closely with product engineering teams to raise the reliability and performance bar across all services.
What You Will Do
Platform & Infrastructure
- Design, deploy, and maintain Kubernetes clusters on both AWS EKS and Azure AKS across multi-region production environments.
- Manage ingress and traffic routing using Traefik Gateway and Nginx Ingress Controllers, including TLS termination and routing rules.
- Implement and maintain Karpenter for intelligent, cost-aware node auto-provisioning on EKS.
- Configure KEDA (Kubernetes Event-Driven Autoscaler) for workload-driven horizontal scaling based on queue depth, custom metrics, and event sources.
- Manage TLS certificate lifecycles across all services using Cert Manager with automated renewal.
- Own Helm chart development, versioning, and dependency management for all platform and application deployments.
Networking & Security
- Design and manage VPC/VNet architectures including NAT Gateways, public/private subnets, route tables, and security groups.
- Configure and maintain Azure Bastion and Hub-Spoke network topologies for secure administrative access.
- Manage secrets, credentials, and certificates using Azure Key Vault and AWS Secrets Manager.
- Administer and secure SFTP server infrastructure for external data exchange partners.
Storage & Messaging
- Manage cloud-native storage solutions including AWS S3 and Azure Blob Storage — lifecycle