Senior Site Reliability Engineer - Cloud Infrastructure (India)

Senior Site Reliability Engineer - Cloud Infrastructure (India)

08 Sep
|
AI Adept Consulting
|
India

08 Sep

AI Adept Consulting

India

Excellent MNC Opportunity - Immediate Joiner only

We're looking for a Senior Site Reliability Engineer to take ownership of reliability, performance, and operational excellence across our cloud infrastructure and Kubernetes platforms. You'll lead the response to complex, high-severity incidents, raise the bar on automation and observability, and mentor other engineers as the team scales.

What You'll Do :

- Lead troubleshooting and resolution of complex, high-severity incidents across AWS infrastructure and Amazon EKS clusters, often serving as incident commander for major incidents.
- Design and build Python-based automation frameworks and tooling that eliminate manual toil and improve reliability at scale, not just one-off scripts.
- Architect and maintain Terraform-based infrastructure-as-code, establishing reusable, secure, and scalable patterns across AWS environments.
- Drive observability strategy - design Grafana dashboards and alerting frameworks that surface the right signals to the right teams at the right time.
- Use SQL and CloudWatch Logs Insights to perform deep root-cause analysis on complex, cross-service incidents.
- Own end-to-end ITSM processes - Incident Management and Problem Management - including root cause analysis, post-incident reviews, and long-term remediation plans.
- Mentor junior and mid-level SREs, reviewing their troubleshooting approach, automation, and incident handling.
- Partner with engineering, product, and leadership teams to communicate incident impact, risk, and remediation plans clearly and confidently.




- Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
- Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
- Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
- Participate in and provide senior-level escalation support for on-call rotations.

What We're Looking For :

- 8 - 10 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, with a track record of owning reliability for production-critical systems.
- Deep, hands-on expertise with core and advanced AWS services (EC2, VPC, IAM, S3, RDS, CloudWatch, networking, etc.).
- Proven expertise troubleshooting complex Amazon EKS issues - cluster-level failures, networking, autoscaling, performance bottlenecks, and upgrade-related issues.
- Strong proficiency in Python for building automation frameworks, internal tooling, and operational systems.




- Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale.
- Strong command of ITSM frameworks, with hands-on ownership of Incident and Problem Management for high-severity issues.
- Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive root-cause analysis.
- Proven experience designing Grafana dashboards and alerting strategies that scale across multiple teams and services.
- Exceptional verbal and written communication skills - able to clearly articulate technical issues, risk, and remediation plans to engineering leadership and non-technical stakeholders alike.
- Experience mentoring or leading other engineers and contributing to team-level reliability strategy.

Mandatory Certifications :

- AWS Certification - required (e.g., AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, or equivalent).
- Certified Kubernetes Administrator (CKA) or equivalent EKS/Kubernetes certification - required.

Soft Skills :

- Calm, decisive leadership during high-pressure, high-severity incidents.
- A strong ownership mindset - drives issues to true resolution and follows through on long-term remediation.
- Natural mentor who raises the technical bar for the team.
- Team-oriented cross-functional partner who works effectively with Dev, Infra, Product, and leadership.

Education :

- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent extensive practical experience.

📌 Senior Site Reliability Engineer - Cloud Infrastructure (India)
🏢 AI Adept Consulting
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer - cloud infrastructure (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer - cloud infrastructure (india) / india