Site Reliability Engineering (SRE) combines software development and systems engineering practices to design, operate, and maintain large-scale, distributed, and highly resilient systems. The primary goal of SRE is to ensure that both internal and customer-facing services consistently deliver high levels of reliability and performance while following sound engineering principles. SRE also focuses on addressing operational challenges through engineering solutions rather than manual intervention.
SRE professionals are accountable for the overall health and operation of systems, leveraging a wide range of tools and methodologies to address operational and technical challenges. Core SRE practices include minimizing time spent on repetitive operational tasks, conducting blameless postmortems, proactively identifying risks, and preventing service disruptions before they occur.
The SRE team thrives on diversity, innovation, curiosity, collaboration, and open communication. Team members bring wide-ranging experiences and perspectives, enabling creative problem-solving and continuous improvement.
We encourage ownership, independent decision-making, and meaningful contributions while providing mentorship and support for ongoing professional development.
Responsibilities
- Maintain system availability and uptime across cloud-native platforms (AWS, GCP) and hybrid environments.
- Develop Infrastructure as Code (IaC) solutions that align with security and engineering standards using technologies such as Terraform, cloud CLI scripting, and cloud SDK programming.
- Design and maintain CI/CD pipelines for application builds, testing, and deployments using Jenkins and cloud-native tooling.
- Create automation tools that facilitate production changes and service request deployments.
- Develop detailed operational runbooks for incident detection, troubleshooting, recovery, and service restoration.
- Analyze and resolve issues across complex distributed architectures and service ecosystem
📌 GCP+ SRE Developer II (Pune)
🏢 UST
📍 Pune