Site Reliability Engineer (Hyderabad)

Site Reliability Engineer (Hyderabad)

20 Aug
|
HCA Healthcare - India
|
Hyderabad

20 Aug

HCA Healthcare - India

Hyderabad

Experience

4+ years

About This Role The Application Support Engineer is a hands-on technical role responsible for the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical applications. The role serves as a senior technical escalation point for complex production issues and drives improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The engineer partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, and Security teams and mentors other support engineers.

Essential Duties

- Serve as a technical escalation point for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting through service restoration.
- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
- Define and improve application reliability and observability practices, including service indicators/objectives, dashboards, alerting, instrumentation, event correlation, and monitoring coverage.
- Conduct application reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop improvement roadmaps.
- Lead toil-reduction efforts and design reusable automation, self-healing, and auto-remediation workflows for repetitive operational activities.
- Guide and support Infrastructure-as-Code, configuration management, GitOps, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
- Support major application upgrades, migrations, platform modernization, patching, setting transitions, and production cutovers.
- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency,



and cost-optimization improvements.
- Provide technical leadership for disaster recovery, resiliency validation, recovery procedures, and rollback readiness.
- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
- Establish and improve technical standards, playbooks, runbooks, templates, operating procedures, and reusable support patterns.
- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, and AI-assisted operations; contribute to proof-of-concept evaluations and recommendations.
- Mentor Application Support Engineers in troubleshooting, monitoring, automation, incident management, root cause analysis, and reliability practices.
- Use incident trends, support metrics, reliability indicators, and operational data to drive measurable continuous improvement.
- Participate in an on-call or senior production escalation rotation where required.

Position Requirements

- 4+ years of experience in Application Support, Production Engineering, Site Reliability Engineering, DevOps, Cloud Operations, or a related technical discipline.
- Strong hands-on experience supporting enterprise or business-critical production applications and troubleshooting across multiple technology layers.
- Advanced experience with monitoring, logging, APM, and observability platforms such as Dynatrace, Splunk, Grafana, or comparable technologies.
- Strong working knowledge of Windows and/or Linux/Unix platforms, relational databases, SQL, APIs, integrations, and distributed application dependencies.
- Experience with cloud platforms such as Azure, GCP, AWS, or comparable enterprise cloud technologies.
- Hands-on experience with CI/CD, source control, GitOps/deployment tools, Infrastructure-as-Code, and automation technologies such as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible.




- Strong understanding of networking concepts including DNS, firewall rules, ports, load balancing, routing, certificates, and application connectivity.
- Strong understanding of identity, privileged access, service accounts, secrets, and enterprise security concepts.
- Demonstrated experience leading complex incident troubleshooting, root cause analysis, and implementation of corrective/preventive improvements.
- Demonstrated ability to mentor engineers, influence cross-functional teams, communicate technical findings, and drive engineering improvements without formal people-management authority.

Education

- Bachelor’s degree in computer science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience may be considered.
- Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.

Knowledge And Skills

Capability Knowledge &

- Skill Expectation

Application &

- Production Engineering

Advanced troubleshooting across complex applications, integrations, dependencies, and production environments. Observability &

- Reliability

Strong knowledge of logs, metrics, traces, dashboards, alerting, service indicators/objectives, performance, availability, and reliability engineering principles.

Cloud &

- Infrastructure

Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies.

DevOps &

- Infrastructure-as-Code

Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices.

Automation &

- Engineering

Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows.

Incident &

- Problem Management

Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements.

Performance, Capacity &

- Security

Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams.

Technical Leadership &

- Improvement

Ability to mentor engineers, lead cross-functional technical resolution, establish standards, evaluate new technologies, and drive measurable improvements.

📌 Site Reliability Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (hyderabad) / hyderabad