Job Title: Site Reliability Engineer
Exp: 5 - 14 Years
Loc: Hyderabad
Work Mode: Hybrid
Duration: 6 Months
:
We're looking for a hands-on Site Reliability Engineer to help keep our cloud infrastructure and Kubernetes workloads reliable, observable, and performant. You'll work closely with engineering and operations teams to troubleshoot production issues, automate repetitive work, and continuously improve how we monitor and respond to incidents.
What You'll Do
- Monitor, troubleshoot, and resolve issues across AWS infrastructure and Amazon EKS clusters, ensuring high availability and performance of containerized workloads.
- Write and maintain Python scripts and tools to automate operational tasks and support troubleshooting workflows.
- Use Terraform to provision, manage, and version-control infrastructure as code across AWS environments.
- Build and maintain Grafana dashboards and alerts that give the team clear visibility into system health, application performance, and infrastructure metrics.
- Query logs and metrics using SQL and CloudWatch Logs Insights to investigate incidents and identify root causes.
- Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
- Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
- Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
- Own the full incident lifecycle detection, triage, communication, resolution, and post-incident review — following ITSM Incident and Problem Management practices.
- Communicate clearly and confidently with technical and non-technical stakeholders during incidents, status updates, and post-incident reviews.
- Contribute to runbooks, knowledge base articles, and process documentation to reduce mean time to resolution (MTTR).
- Participate in on-call rotations and respond promptly to production alerts.
What We're Looking For
- 4 to 6 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
- Strong working knowledge of core AWS services (EC2, VPC, IAM, S3, RDS, CloudWatch, etc.).
- Hands-on experience troubleshooting Amazon EKS — pod failures, networking issues, node health, scaling problems, and workload performance.
- Proficiency in Python for scripting, automation, and internal tooling.
- Practical experience writing and managing Terraform modules for AWS infrastructure.
- Experience working within an ITSM framework, specifically Incident Management and Problem Management.
- Comfortable querying data with SQL and AWS CloudWatch Logs Insights to investigate issues.
- Experience building dashboards and configuring alerts in Grafana or a similar observability tool.
- Strong verbal and written communication skills, with the ability to clearly explain technical issues, root causes, and remediation steps to varied audiences.
Mandatory to Have
- AWS certification (e.g., AWS Certified Solutions Architect – Associate, AWS Certified SysOps Administrator, or AWS Certified DevOps Engineer).
- Certified Kubernetes Administrator (CKA) or an equivalent EKS/Kubernetes certification.
Soft Skills
- A calm, structured approach to troubleshooting under pressure.
- An ownership mindset — follows through on issues until they're truly resolved.
- A team-oriented team player who works well across Dev, Infra, and Product teams
If interested, please share your resume to the mail ID below:
[email protected]
📌 Site Reliability Engineer (Hyderabad)
🏢 CIEL HR
📍 Hyderabad