Chennai, Tamil Nadu
Job Summary
We are seeking a highly motivated Site Reliability Engineer (SRE) responsible for ensuring the reliability, availability, scalability, and performance of enterprise applications and infrastructure. The ideal candidate will combine software engineering, automation, cloud operations, and observability expertise to reduce operational toil and improve platform resilience through automation and proactive engineering practices.
Key Responsibilities
Ensure high availability, performance, and scalability of critical business services.Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).Lead major incident response, root cause analysis (RCA), and post-incident reviews.Implement proactive monitoring, alerting, and self-healing capabilities.Conduct capacity planning, performance tuning, and reliability assessments.Automation & EngineeringDesign and implement Infrastructure as Code (IaC) solutions using Terraform, Ansible, or similar tools.Develop automation scripts using Python, PowerShell, Bash, or equivalent technologies.Drive CI/CD pipeline enhancements and deployment automation.Reduce manual effort (toil) through engineering-led automation initiatives. Cloud & Platform Management Manage and optimize Azure, AWS, GCP,
and hybrid-cloud environments.Support containerized workloads using Kubernetes and Docker.Implement high-availability, disaster recovery, and resilience architectures.Collaborate with infrastructure, application, security, and DevOps teams to improve platform stability. Observability & Monitoring Implement and maintain observability platforms for logs, metrics, traces, and events.Establish monitoring dashboards, reliability KPIs, and performance baselines.Drive noise reduction, event correlation, and predictive analytics initiatives.Utilize tools such as Azure Monitor, Dynatrace, AppDynamics, ELK, Grafana, Prometheus, Splunk, or equivalent
Skill Requirements
Solid experience in Cloud Platforms (Azure/AWS/GCP).Expertise in Linux and/or Windows administration.Hands-on experience with Terraform, Ansible, Puppet, or similar automation tools.Strong scripting/programming skills in Python, PowerShell, Bash, or Go.Experience with CI/CD tools such as Azure DevOps, Jenkins, GitHub Actions, or GitLab.Good understanding of networking, storage, databases, and middleware technologies.Experience with container orchestration platforms such as Kubernetes
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 Sr Subject Matter Expert (Support&Ops) (India)
🏢 HCLTech
📍 India