Senior Consultant - Site Reliability Engineer (Hyderabad)

Senior Consultant - Site Reliability Engineer (Hyderabad)

07 Oct
|
HCA Healthcare - India
|
Hyderabad

07 Oct

HCA Healthcare - India

Hyderabad

General Position Information

Reports Directly To (Title)

Manager –

- Site Reliability Engineer

Matrix Reports To (Title)

As applicable based on functional alignment

Direct Reports

Individual contributor; no direct reports

Position Summary The Senior Consultant - Site Reliability Engineering (SRE) is a hands-on technical leadership and consulting role responsible for driving the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical services. The role serves as a principal technical escalation point and trusted advisor for complex production challenges, shaping SRE strategy and engineering improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness.

The Senior

Consultant partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, Security, and business stakeholders; leads cross-functional initiatives; and mentors senior engineers.

Responsibilities

- Serve as a principal technical escalation point and SRE consultant for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting, executive-level technical communication, and service restoration.
- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
- Define, govern, and mature SRE and observability practices, including SLIs/SLOs, error budgets, dashboards, alerting, instrumentation, event correlation, monitoring coverage, and reliability reporting.
- Lead enterprise reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop prioritized multi-quarter improvement roadmaps.
- Lead strategic toil-reduction programs and design reusable automation, self-healing, and auto-remediation patterns that improve engineering productivity and service reliability at scale.
- Provide technical governance and hands-on leadership for Infrastructure-as-Code, configuration management, GitOps, platform engineering, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
- Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers.
- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.
- Own technical direction for disaster recovery, business continuity engineering, resiliency validation, recovery objectives/procedures, failure testing, and rollback readiness.
- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
- Establish SRE standards, reference architectures, playbooks, runbooks, templates,



operating procedures, governance mechanisms, and reusable engineering patterns across teams.
- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, platform engineering, and AI-assisted operations; lead proof-of-concept evaluations, technical recommendations, and adoption roadmaps.
- Mentor senior SRE, Production Engineering, DevOps, and Application Support engineers; provide technical coaching in troubleshooting, observability, automation, incident management, root cause analysis, architecture, and reliability practices.
- Use incident trends, operational metrics, SLO performance, risk indicators, and engineering data to define reliability priorities, influence stakeholders, and drive measurable continuous improvement.
- Provide consultative leadership to application and platform teams on reliability architecture, SRE adoption, production readiness, cloud modernization, and operational risk reduction.
- Lead reliability reviews with senior stakeholders, translate technical risk into business impact, and define measurable remediation plans, success criteria, and governance checkpoints.
- Participate in an on-call or senior production escalation rotation where required.

Education &

- Experience

- Application &
- Production Engineering - Advanced troubleshooting across complex applications, integrations, dependencies, and production environments.
- Observability &
- Reliability - Advanced proficiency in Excel and presentation tools; comfortable working with large data sets, pivots, lookups, and structured trackers.
- Cloud &
- Infrastructure - Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies.
- DevOps &
- Infrastructure-as-Code - Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices.
- Automation &
- Engineering - Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows.
- Incident &
- Problem Management - Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements.
- Performance, Capacity &
- Security - Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams.
- Technical Leadership &
- Improvement - Expert ability to operate as a senior SRE consultant, mentor experienced engineers, influence architecture and stakeholders, establish enterprise standards, lead transformation roadmaps, and drive measurable reliability improvements.

Must Have Skills

- Advanced expertise in Site Reliability Engineering (SRE), Production Support, and Application Operations, including troubleshooting complex enterprise applications and distributed systems.
- Strong hands-on experience with Observability and Monitoring platforms such as Dynatrace, Splunk, Grafana, Prometheus, OpenTelemetry, logging, tracing,



alerting, and telemetry analytics.
- Deep knowledge of Cloud Platforms and Infrastructure Technologies including Azure, AWS, GCP, Linux/Unix, Windows, networking, databases, and enterprise architecture dependencies.
- Extensive experience with DevOps, CI/CD, GitOps, and Infrastructure-as-Code using tools such as Terraform, Ansible, Azure DevOps, GitHub, GitLab, Argo CD, and related automation frameworks.
- Proven ability to lead major incident management, root cause analysis (RCA), problem management, and service restoration efforts for business-critical applications and platforms.
- Strong expertise in automation engineering, scripting, self-healing solutions, operational tooling, and reducing manual operational effort through scalable automation.
- Advanced understanding of Performance Engineering, Capacity Planning, Disaster Recovery, Business Continuity, Security, Identity Management, and Operational Resilience principles.
- Demonstrated experience defining and implementing SLIs, SLOs, Error Budgets, Reliability Metrics, and Production Readiness standards to improve service reliability and operational excellence.
- Ability to provide technical leadership, mentoring, and consultative guidance to engineering teams, influencing architectural decisions and driving reliability-focused transformation initiatives.
- Excellent communication, stakeholder management, and problem-solving skills with the ability to translate technical risks into business impacts and drive measurable improvement outcomes.

Nice To Have Skills

- Experience implementing and governing SRE practices such as SLIs, SLOs, Error Budgets, Reliability Reviews, and Production Readiness Assessments.
- Hands-on experience with Kubernetes, container platforms, service mesh technologies, and cloud-native application architectures.
- Experience with AIOps, predictive analytics, machine learning-assisted monitoring, automated remediation, and intelligent operations platforms.
- Expertise in large-scale cloud modernization, platform engineering, and migration initiatives across Azure, AWS, or Google Cloud environments.
- Relevant certifications such as Google Professional Cloud Engineer, Azure Solutions Architect, Certified Kubernetes Administrator (CKA), Terraform Associate, ITIL, or related SRE/DevOps certifications.

Licenses, Certifications &

- Training

- N/A

Knowledge, Skills, Abilities, Behaviors

- Strong leadership, consulting, and stakeholder management skills with the ability to influence technical strategy, drive reliability initiatives, and communicate complex technical concepts to engineering and business leaders.
- Exceptional analytical, troubleshooting, and problem-solving abilities with expertise in diagnosing complex production issues, assessing technical risks, and implementing sustainable solutions.
- Demonstrated ability to mentor and guide senior engineers, foster collaboration across teams, and promote Site Reliability Engineering best practices and operational excellence.
- Ability to operate effectively in high-pressure environments, manage competing priorities, and make sound decisions during critical incidents while maintaining a focus on business outcomes.
- Solid ownership mindset, adaptability, and commitment to continuous improvement, with a focus on automation, innovation, reliability, customer experience, and long-term operational maturity.

📌 Senior Consultant - Site Reliability Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior consultant - site reliability engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: senior consultant - site reliability engineer (hyderabad) / hyderabad