We are looking for candidates with robust experience in production support, AWS cloud technologies, automation, and reliability engineering across both on-premises and cloud environments.
LONGDESCRIPTION section. 2 of 6.Section Title: Key Responsibilities
Key Responsibilities
Key Responsibilities
- Own end-to-end availability, reliability, and performance of critical production applications.
- Monitor system health against SLOs and proactively prevent service disruptions.
- Lead incident management, troubleshooting, and post-incident reviews.
- Provide on-call support and drive timely resolution of critical incidents.
- Build and maintain observability through monitoring, logging, dashboards, and alerting.
- Own CI/CD pipelines and ensure safe, automated deployments.
- Implement Infrastructure as Code (IaC) and drive operational automation.
- Ensure platform security, compliance, and audit readiness.
- Drive capacity planning, performance optimisation, resilience testing, and disaster recovery readiness.
- Collaborate with engineering teams to continuously improve platform stability and reliability.
📌 SRE +Splunk - 6+ years - Chennai, Bangalore, Noida
🏢 HCLTech
📍 Noida