Jon Title:-Senior Site Reliability Engineer (SRE) / DevOps Engineer
Location: Pune - In Office
Experience: 10+ plus years of experience
On-Call Rotation Required (24/7 Production Support)
About the Role
We are seeking a Senior Site Reliability Engineer (SRE) / DevOps Engineer who will be responsible for ensuring the reliability, scalability, security, and performance of our production systems across multi-cloud environments (AWS, GCP, Azure). This role combines solid DevOps automation expertise with true SRE ownership — including on-call participation, incident management, root cause analysis, reliability engineering, and proactive system improvements. The ideal candidate balances incident response and firefighting with long-term engineering improvements that reduce toil, improve SLAs, and strengthen system resilience.
Key Responsibilities
1.Incident Response & On-Call Ownership
Participate in 24/7 on-call rotation for production systems
Rapidly diagnose, mitigate, and resolve high-severity incidents
Lead Root Cause Analysis (RCA) and post-mortem documentation
Implement corrective and preventive measures to avoid recurrence
Maintain SLAs/SLOs and reduce Mean Time to Recovery (MTTR)
2.Reliability Engineering & System Hardening
Design and implement reliability improvements to increase availability and reduce system fragility
Engineer solutions to eliminate repetitive operational work (toil reduction)
Improve redundancy, failover strategies, and disaster recovery planning
Track and improve SRE metrics (availability, latency, error rates, capacity)
3. Infrastructure & Cloud Engineering (Multi-Cloud)
Manage and optimize infrastructure across:
AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)
Google Cloud Platform (GCP) (Compute Engine, Cloud Storage, Cloud SQL, IAM, VPC) (Having GCP is a plus)
Microsoft Azure (Virtual Machines, Networking, Storage, Azure Monitor)
Administer and optimize Kubernetes clusters
Manage Helm deployments and containerized workloa
📌 Senior Site Reliability Engineer (Pune)
🏢 AcquireX
📍 Pune