We are looking for a skilled Site Reliability Engineer (SRE) to improve the availability, performance, scalability, and operational resilience of our production services. The role focuses on building reliable systems, reducing operational complexity, and ensuring that business-critical applications remain stable, efficient, and highly available.
The candidate will be responsible for defining and monitoring Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to measure service reliability and establish clear performance expectations. You will design and maintain monitoring, alerting, and observability solutions to provide visibility into application and infrastructure health.
The role involves automating repetitive operational tasks and reducing manual intervention through scripting, infrastructure automation, and reliable deployment practices. You will investigate production incidents, troubleshoot system and application issues, coordinate recovery activities, and conduct blameless post-incident reviews to identify root causes and prevent recurring problems.
Strong knowledge of Linux, networking, scripting, distributed systems, cloud platforms, Kubernetes, and Terraform is required. Experience with observability and monitoring tools such as Prometheus, Grafana, ELK, OpenTelemetry,
or similar technologies is expected. The candidate should also have a valuable understanding of CI/CD, incident response, capacity planning, disaster recovery, and reliability engineering practices.
You will work closely with software developers, DevOps engineers, cloud teams, and other technical stakeholders to design reliable services, improve deployment safety, and identify potential reliability risks throughout the application lifecycle. The role also includes supporting production releases, improving system performance, and ensuring that infrastructure can handle changing workloads and business requirements.
The ideal candidate should have strong problem-solving and troubleshooting skills, an automation-focused mindset, and the ability to work effectively during production incidents. Programming experience with Python, Go, Java, or a comparable language is required to build automation tools and improve operational workflows.
You will also contribute to capacity planning, disaster-recovery strategies, system resilience, and continuous reliability improvements, while maintaining clear operational documentation, runbooks, monitoring standards, and incident-management procedures.
📌 Site Reliability Engineer (SRE) (Chennai)
🏢 C1X
📍 Chennai