Were looking for a Senior Site Reliability Engineer to take ownership of reliability performance and operational excellence across our cloud infrastructure and Kubernetes platforms Youll lead the response to complex highseverity incidents raise the bar on automation and observability and mentor other engineers as the team scales
What Youll Do
Lead troubleshooting and resolution of complex highseverity incidents across AWS infrastructure and Amazon EKS clusters often serving as incident commander for major incidents
Design and build Pythonbased automation frameworks and tooling that eliminate manual toil and improve reliability at scale not just oneoff scripts
Architect and maintain Terraformbased infrastructureascode establishing reusable secure and scalable patterns across AWS environments
Drive observability strategy design Grafana dashboards and ing frameworks that surface the right signals to the right teams at the right time
Use SQL and CloudWatch Logs Insights to perform deep rootcause analysis on complex crossservice incidents
Own endtoend ITSM processes Incident Management and Problem Management including root cause analysis postincident reviews and longterm remediation plans
Mentor junior and midlevel SREs reviewing their troubleshooting approach automation and incident handling
Partner with engineering product and leadership teams to communicate incident impact risk and remediation plans clearly and confidently
Design and build selfhealing automation and runbooks that detect known failure patterns and trigger remediation automatically reducing manual intervention and recovery time for recurring incidents
Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health latency and failover readiness across all deployment zones
Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks
Participate in and provide seniorlevel escalation support for oncall rotations
What Were Looking For
8 to 10 years of experience in Site Reliability Engineering DevOps or Cloud Infrastructure roles with a track record of owning reliability for productioncritical systems
Deep handson expertise with core and advanced AWS services EC2 VPC IAM S3 RDS CloudWatch networking etc
Strong proficiency in Python for building automation frameworks internal tooling and operational systems
Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale
Strong command of ITSM frameworks with handson ownership of Incident and Problem Management for highseverity issues
Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive rootcause analysis
Proven experience designing Grafana dashboards and ing strategies that scale across multiple teams and services
Exceptional verbal and written communication skills able to clearly articulate technical issues risk and remediation plans to engineering leadership and nontechnical stakeholders alike
Experience mentoring or leading other engineers and contributing to teamlevel reliability strategy
Mandatory Certifications
- AWS Certification required eg AWS Certified Solutions Architect Professional AWS Certified DevOps Engineer Skilled or equivalent
- Certified Kubernetes Administrator CKA or equivalent EKSKubernetes certification required
Soft Skills Calm decisive leadership during highpressure highseverity incidents A strong ownership mindset drives issues to true resolution and follows through on longterm remediation Natural mentor who raises the technical bar for the team
Collaborative crossfunctional partner who works effectively with Dev Infra Product and leadership