We are looking for candidates with robust experience in production support, AWS cloud technologies, automation, and reliability engineering across both on-premises and cloud environments.
LONGDESCRIPTION section. 2 of 6.Section Title: Key Responsibilities
Key Responsibilities
Key Responsibilities
Own end-to-end availability, reliability, and performance of critical production applications.
Monitor system health against SLOs and proactively prevent service disruptions.
Lead incident management, troubleshooting, and post-incident reviews.
Provide on-call support and drive timely resolution of critical incidents.
Build and maintain observability through monitoring, logging, dashboards, and alerting.
Own CI/CD pipelines and ensure secure, automated deployments.
Implement Infrastructure as Code (IaC) and drive operational automation.
Ensure platform security, compliance, and audit readiness.
Drive capacity planning, performance optimisation, resilience testing, and disaster recovery readiness.
Collaborate with engineering teams to continuously improve platform stability and reliability.
📌 Sre +splunk 6+ Years Chennai, Bangalore, Noida
🏢 HCLTech
📍 Noida