Lead Engineer – Site Reliability Engineering
Role Purpose : Ensures high availability, performance, scalability, and resilience of cloud and infrastructure platforms by applying SRE engineering principles, automation-first practices, observability, and continual reliability improvements across services and platforms.
Key Responsibilities:
· Implement SRE frameworks, SLIs/SLOs/SLAs, error budgets, performance engineering, and reliability guardrails across cloud platforms & services.
· Drive automation for provisioning, deployment, configuration management, drift control, patching, recovery, and operations workflows.
· Build observability stack, dashboards, anomaly detection, synthetic tests, runbooks, incident readiness and RCA automation.
· Partner with DevOps,
Platform Engineering, Cloud Engineering & application squads to define reliability patterns, capacity planning & scalable workload landing models.
· Lead incident response, major incident coordination, postmortem improvement actions, resiliency testing, fault injection, chaos engineering initiatives.
· Ensure infra security alignment, vulnerability remediation, compliance & secure configuration baselines in cloud infrastructure.
· Mentor engineers in SRE best practices, automation pipelines, tooling standardization and operational excellence.
📌 Lead Engineer – Site Reliability Engineering (Chennai)
🏢 CBTS
📍 Chennai