Job Summary
Site Reliability Engineer - AI Operations to implement Site Reliability Engineering practices with a strong focus on automation, observability, AIOps, Gen AI, and self-healing operations. This role will help improve production reliability, reduce manual toil, accelerate incident response, and enable intelligent operational automation across enterprise applications and platforms.
The ideal candidate combines solid production operations experience with engineering, automation, observability, and AI-enabled operations capabilities. This role will work closely with application, infrastructure, cloud, security, operations, and leadership teams to strengthen reliability, improve operational efficiency, and modernize how production services are supported.
Location: Hyderabad, India
Key Responsibilities
- Site Reliability Engineering & Operations
- Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.
- Support production reliability by improving availability, performance, scalability, and operational readiness of applications and platforms.
- Participate in incident response, troubleshooting, root cause analysis, postmortems, and corrective/preventive action planning.
- Drive shift-left reliability by embedding monitoring, alerting, automation,
and operational readiness into SDLC and CI/CD pipelines.
- Support reliable application and platform operations in Azure Cloud.
- Automation & Self-Healing
- Build and maintain automation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR.
- Identify repetitive operational activities and automate them where appropriate.
- Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.
- Create and maintain reusable automation playbooks, operational procedures, and knowledge articles.
- Ensure automation activities follow change management, governance, security, and compliance processes.
- Observability, Monitoring & Reliability Metrics
Observability, Monitoring & Reliability Metrics
- Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.
- Work with monitoring and APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.
Disclaimer: This job description has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.
📌 SDM SRE Senior Engineer (Hyderabad)
🏢 TJX
📍 Hyderabad