Hello All,
Greetings from ZettaMine!!!
Role: Senior Site Reliability Engineer (SRE) – AWS & EKS
Experience: 4 Years & 7 Years
Location: Bangalore, Pune, Chennai
Notice Period: Immediate Joiners
Shift Timing: 2:00 PM – 10:00 PM
Role: Senior Site Reliability Engineer
Primary Skills: AWS, Amazon EKS, Kubernetes, Python, Terraform
Mandatory Certifications: AWS Certification CKA/Kubernetes Certification
Hiring for one of the Big 4's Clients
We are looking for Immediate to 15 Days Joiners only
Job Summary
We are looking for an experienced Senior Site Reliability Engineer (SRE) with deep hands-on expertise in AWS cloud infrastructure, Amazon EKS/Kubernetes, Python automation, Terraform, observability, and production incident management.
The ideal candidate will take ownership of reliability, availability, performance, and operational excellence across cloud infrastructure and Kubernetes platforms. This role requires strong expertise in resolving complex production incidents, building scalable automation and self-healing solutions, improving observability, and mentoring engineers.
Key Responsibilities
- Lead troubleshooting and resolution of complex, high-severity production incidents across AWS infrastructure and Amazon EKS clusters.
- Act as an Incident Commander for major incidents and drive issues through complete resolution.
- Troubleshoot complex EKS/Kubernetes issues, including cluster-level failures, networking issues, autoscaling problems, performance bottlenecks, and upgrade-related incidents.
- Design and develop Python-based automation frameworks, internal tools, and operational solutions to eliminate manual effort and improve platform reliability.
- Architect, implement, and maintain Terraform-based Infrastructure as Code (IaC) with reusable, secure,
and scalable modules and patterns.
- Design and implement self-healing automation and automated runbooks to detect known failure patterns and trigger remediation automatically.
- Drive the observability strategy by developing Grafana dashboards, monitoring solutions, and scalable alerting frameworks.
- Implement and maintain multi-region monitoring to provide consistent visibility into system health, latency, availability, and failover readiness.
- Perform advanced troubleshooting and root-cause analysis (RCA) using SQL and AWS CloudWatch Logs Insights.
- Own end-to-end ITSM Incident Management and Problem Management, including RCA, post-incident reviews, corrective actions, and long-term remediation.
- Proactively identify potential failure points, reliability risks, and performance bottlenecks before they impact production.
- Reduce operational workload by automating repetitive and recurring manual activities.
- Provide senior-level escalation support and participate in on-call rotations.
- Mentor junior and mid-level SRE/DevOps engineers in troubleshooting, automation, incident handling, and reliability engineering best practices.
- Partner with engineering, product, and leadership teams to clearly communicate incident impact, technical risks, root causes, and remediation plans.
Mandatory Certifications
Candidates must have both of the following:
1. AWS Certification, such as:
- AWS Certified Solutions Architect – Skilled
- AWS Certified DevOps Engineer – Professional
- Or equivalent advanced AWS certification
1. Kubernetes Certification
- Certified Kubernetes Administrator (CKA)
- Or equivalent EKS/Kubernetes certification
Interested candidates share your updated CV to
[email protected] or WhatsApp to (phone hidden).
Thanks& Regards,
Praneeth
📌 Senior Site Reliability Engineer (SRE) – AWS & EKS (Bengaluru)
🏢 ZettaMine Labs
📍 Bengaluru