13 Aug
|
Apptad
|
Hyderabad
- Work Location Bangalore/Hyderabad/Pune
- Work Mode (Onsite / Hybrid / Remote) – Hybrid
Senior Site Reliability Engineer / Platform Engineer – AWS EKS
Role Overview
This role owns the reliability, security, automation, and scalability of the AWS EKS platform. The Senior SRE / Platform Engineer acts as a platform owner, enabling application teams through standardized infrastructure, GitOps workflows, policy enforcement, and observability. The position requires deep hands-on expertise across Kubernetes, AWS infrastructure, security tooling, and SRE operational practices.
Primary Responsibilities
Kubernetes & EKS Platform Engineering
- Architect, deploy, and operate production-grade Kubernetes platforms on AWS EKS
- Implement EKS automation using EKS Blueprints and manage lifecycle of all EKS add-ons
- Plan and execute Kubernetes and EKS version upgrades with minimal downtime
Autoscaling & Compute Optimization
- Design and operate Karpenter-based autoscaling for stateless and stateful workloads
- Optimize cost, performance, and availability through dynamic instance provisioning
Service Mesh & Traffic Management
- Design and operate Istio service mesh (sidecar and ambient mesh models)
- Implement traffic policies including mTLS, retries, circuit breaking, and timeouts
Policy, Security & Runtime Protection
- Implement Kubernetes admission and governance using Kyverno and OPA/Gatekeeper
- Operate Falco for runtime threat detection and incident investigation
- Integrate security controls into GitOps workflows
Infrastructure as Code & Automation
- Develop reusable Terraform modules for AWS infrastructure including VPCs, EKS, and Transit Gateway
- Implement Terragrunt-based multi-account, multi-region architectures
GitOps & Platform Operations
- Design and operate self-managed Argo CD for platform, security, and add-on management
- Define Git-based promotion and access models across environments
Observability & SRE Operations
- Design and operate Prometheus-based monitoring and alerting
- Participate in incident response, root cause analysis, and reliability improvements
- Reduce operational toil through automation and self-service tooling
Security & Compliance (Wiz)
- Own remediation of Wiz-reported security findings across AWS infrastructure and EKS clusters
- Partner with Security teams to implement preventive guardrails
Required Skills & Experience
- 6+ years of experience in SRE, Platform Engineering, or Cloud Infrastructure roles
- Strong hands-on expertise with Kubernetes and AWS EKS
- Deep experience with Terraform, Terragrunt, and EKS Blueprints
- Hands-on experience with Karpenter autoscaling
- Experience operating Istio service mesh
- Strong experience with Kyverno, OPA/Gatekeeper, and Falco
- Observability experience using Prometheus
- GitOps experience using Argo CD
- Robust AWS fundamentals: VPC, IAM, EC2, ALB/NLB, EBS/EFS
- Proven experience resolving Wiz security findings
Nice to Have
- Experience with Ambient Mesh
- Familiarity with SLOs, SLIs, and error budgets
- Experience in large-scale or regulated environments
📌 SRE, Platform Engineer - AWS EKS For Apptad Technology _Hybrid (Hyderabad)
🏢 Apptad
📍 Hyderabad