09 Sep
|
HCL Tech
|
India
Job Summary
Role: Site Reliability Engineer (SRE)
Experience : 6 to 12 Years
Location: Chennai, Bangalore, Gurgaon
We are looking for candidates with solid experience in production support, AWS cloud technologies, automation, and reliability engineering across both on-premises and cloud settings.
Key Responsibilities
Key Responsibilities
Own end-to-end availability, reliability, and performance of critical production applications.
Monitor system health against SLOs and proactively prevent service disruptions.
Lead incident management, troubleshooting, and post-incident reviews.
Provide on-call support and drive timely resolution of critical incidents.
Build and maintain observability through monitoring, logging, dashboards, and alerting.
Own CI/CD pipelines and ensure protected, automated deployments.
Implement Infrastructure as Code (IaC) and drive operational automation.
Ensure platform security, compliance, and audit readiness.
Drive capacity planning, performance optimisation, resilience testing, and disaster recovery readiness.
Collaborate with engineering teams to continuously improve platform stability and reliability.
Skill Requirements
Required Skills
AWS Cloud (EC2, Lambda, S3, and related services)
Terraform / Infrastructure as Code (IaC)
CI/CD tooling (Jenkins, GitHub Actions, GitLab CI/CD, AWS DevOps tools)
Monitoring & Observability (Grafana, Splunk, Prometheus, Datadog)
Incident, Problem, and Change Management
Linux/Unix Administration
Python, Bash, or PowerShell scripting
Production Support and Reliability Engineering
Other Requirements
📌 Senior Technical Lead Chennai (India)
🏢 HCL Tech
📍 India