08 Sep
|
HCL Tech
|
Chennai
Job Summary
Role: Site Reliability Engineer (SRE)
Experience : 6 to 12 Years
Location: Chennai, Bangalore, Gurgaon
We are looking for candidates with solid experience in production support, AWS cloud technologies, automation, and reliability engineering across both on-premises and cloud environments.
Key Responsibilities
Key Responsibilities
- Own end-to-end availability, reliability, and performance of critical production applications.
- Monitor system health against SLOs and proactively prevent service disruptions.
- Lead incident management, troubleshooting, and post-incident reviews.
- Provide on-call support and drive timely resolution of critical incidents.
- Build and maintain observability through monitoring, logging, dashboards, and alerting.
- Own CI/CD pipelines and ensure safe, automated deployments.
- Implement Infrastructure as Code (IaC) and drive operational automation.
- Ensure platform security, compliance, and audit readiness.
- Drive capacity planning, performance optimisation, resilience testing, and disaster recovery readiness.
- Collaborate with engineering teams to continuously improve platform stability and reliability.
Skill Requirements
Required Skills
- AWS Cloud (EC2, Lambda, S3, and related services)
- Terraform / Infrastructure as Code (IaC)
- CI/CD tooling (Jenkins, GitHub Actions, GitLab CI/CD, AWS DevOps tools)
- Monitoring & Observability (Grafana, Splunk, Prometheus, Datadog)
- Incident, Problem, and Change Management
- Linux/Unix Administration
- Python, Bash, or PowerShell scripting
- Production Support and Reliability Engineering
Other Requirements
📌 Senior Technical Lead (Chennai)
🏢 HCL Tech
📍 Chennai