Site Reliability Engineer (Hyderabad)

Site Reliability Engineer (Hyderabad)

02 Oct
|
HCA Healthcare UK
|
Hyderabad

02 Oct

HCA Healthcare UK

Hyderabad

Level II - Site Reliability Engineer | HCA Healthcare

Job Summary

The Application Support Engineer is a hands-on technical role responsible for the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical applications. The role serves as a senior technical escalation point for complex production issues and drives improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The engineer partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, and Security teams and mentors other support engineers.

General Position Information

- Reports directly to (Title): Manager - Platform Engineer

- Matrix reports to (Title): As applicable based on functional alignment

- Direct Reports: Individual contributor; no direct reports

Responsibilities

- Serve as a technical escalation point for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting through service restoration.

- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.

- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.

- Define and improve application reliability and observability practices, including service indicators/objectives, dashboards, alerting,



instrumentation, event correlation, and monitoring coverage.

- Conduct application reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop improvement roadmaps.

- Lead toil-reduction efforts and design reusable automation, self-healing, and auto-remediation workflows for repetitive operational activities.

- Guide and support Infrastructure-as-Code, configuration management, GitOps, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.

- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.

- Support major application upgrades, migrations, platform modernization, patching, setting transitions, and production cutovers.

- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.

- Provide technical leadership for disaster recovery, resiliency validation, recovery procedures, and rollback readiness.





- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.

- Establish and improve technical standards, playbooks, runbooks, templates, operating procedures, and reusable support patterns.

- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, and AI-assisted operations; contribute to proof-of-concept evaluations and recommendations.

- Mentor Application Support Engineers in troubleshooting, monitoring, automation, incident management, root cause analysis, and reliability practices.

- Use incident trends, support metrics, reliability indicators, and operational data to drive measurable continuous improvement.

- Participate in an on-call or senior production escalation rotation where required.

Education and Experience

- Bachelor's degree in computer science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience may be considered.

- 4+ years of experience in Application Support, Production Engineering, Site Reliability Engineering, DevOps, Cloud Operations, or a related technical discipline.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Site Reliability Engineer (Hyderabad)
🏢 HCA Healthcare UK
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (hyderabad) / hyderabad