07 Oct
|
HCA Healthcare - India
|
Hyderabad
07 Oct
HCA Healthcare - India
Hyderabad
General Position Information
Reports Directly To (Title)
Manager –
- Site Reliability Engineer
Matrix Reports To (Title)
As applicable based on functional alignment
Direct Reports
Individual contributor; no direct reports
Position Summary The Senior Consultant - Site Reliability Engineering (SRE) is a hands-on technical leadership and consulting role responsible for driving the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical services. The role serves as a principal technical escalation point and trusted advisor for complex production challenges, shaping SRE strategy and engineering improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness.
The Senior
Consultant partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, Security, and business stakeholders; leads cross-functional initiatives; and mentors senior engineers.
Responsibilities
- Serve as a principal technical escalation point and SRE consultant for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting, executive-level technical communication, and service restoration.
- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
- Define, govern, and mature SRE and observability practices, including SLIs/SLOs, error budgets, dashboards, alerting, instrumentation, event correlation, monitoring coverage, and reliability reporting.
- Lead enterprise reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop prioritized multi-quarter improvement roadmaps.
- Lead strategic toil-reduction programs and design reusable automation, self-healing, and auto-remediation patterns that improve engineering productivity and service reliability at scale.
- Provide technical governance and hands-on leadership for Infrastructure-as-Code, configuration management, GitOps, platform engineering, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
- Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers.
- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.
- Own technical direction for disaster recovery, business continuity engineering, resiliency validation, recovery objectives/procedures, failure testing, and rollback readiness.
- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
- Establish SRE standards, reference architectures, playbooks, runbooks, templates,
operating procedures, governance mechanisms, and reusable engineering patterns across teams.
- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, platform engineering, and AI-assisted operations; lead proof-of-concept evaluations, technical recommendations, and adoption roadmaps.
- Mentor senior SRE, Production Engineering, DevOps, and Application Support engineers; provide technical coaching in troubleshooting, observability, automation, incident management, root cause analysis, architecture, and reliability practices.
- Use incident trends, operational metrics, SLO performance, risk indicators, and engineering data to define reliability priorities, influence stakeholders, and drive measurable continuous improvement.
- Provide consultative leadership to application and platform teams on reliability architecture, SRE adoption, production readiness, cloud modernization, and operational risk reduction.
- Lead reliability reviews with senior stakeholders, translate technical risk into business impact, and define measurable remediation plans, success criteria, and governance checkpoints.
- Participate in an on-call or senior production escalation rotation where required.
Education &
- Experience
- Application &
- Production Engineering - Advanced troubleshooting across complex applications, integrations, dependencies, and production environments.
- Observability &
- Reliability - Advanced proficiency in Excel and presentation tools; comfortable working with large data sets, pivots, lookups, and structured trackers.
- Cloud &
- Infrastructure - Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies.
- DevOps &
- Infrastructure-as-Code - Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices.
- Automation &
- Engineering - Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows.
- Incident &
- Problem Management - Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements.
- Performance, Capacity &
- Security - Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams.
- Technical Leadership &
- Improvement - Expert ability to operate as a senior SRE consultant, mentor experienced engineers, influence architecture and stakeholders, establish enterprise standards, lead transformation roadmaps, and drive measurable reliability improvements.
Must Have Skills
- Advanced expertise in Site Reliability Engineering (SRE), Production Support, and Application Operations, including troubleshooting complex enterprise applications and distributed systems.
- Strong hands-on experience with Observability and Monitoring platforms such as Dynatrace, Splunk, Grafana, Prometheus, OpenTelemetry, logging, tracing,
alerting, and telemetry analytics.
- Deep knowledge of Cloud Platforms and Infrastructure Technologies including Azure, AWS, GCP, Linux/Unix, Windows, networking, databases, and enterprise architecture dependencies.
- Extensive experience with DevOps, CI/CD, GitOps, and Infrastructure-as-Code using tools such as Terraform, Ansible, Azure DevOps, GitHub, GitLab, Argo CD, and related automation frameworks.
- Proven ability to lead major incident management, root cause analysis (RCA), problem management, and service restoration efforts for business-critical applications and platforms.
- Strong expertise in automation engineering, scripting, self-healing solutions, operational tooling, and reducing manual operational effort through scalable automation.
- Advanced understanding of Performance Engineering, Capacity Planning, Disaster Recovery, Business Continuity, Security, Identity Management, and Operational Resilience principles.
- Demonstrated experience defining and implementing SLIs, SLOs, Error Budgets, Reliability Metrics, and Production Readiness standards to improve service reliability and operational excellence.
- Ability to provide technical leadership, mentoring, and consultative guidance to engineering teams, influencing architectural decisions and driving reliability-focused transformation initiatives.
- Excellent communication, stakeholder management, and problem-solving skills with the ability to translate technical risks into business impacts and drive measurable improvement outcomes.
Nice To Have Skills
- Experience implementing and governing SRE practices such as SLIs, SLOs, Error Budgets, Reliability Reviews, and Production Readiness Assessments.
- Hands-on experience with Kubernetes, container platforms, service mesh technologies, and cloud-native application architectures.
- Experience with AIOps, predictive analytics, machine learning-assisted monitoring, automated remediation, and intelligent operations platforms.
- Expertise in large-scale cloud modernization, platform engineering, and migration initiatives across Azure, AWS, or Google Cloud environments.
- Relevant certifications such as Google Professional Cloud Engineer, Azure Solutions Architect, Certified Kubernetes Administrator (CKA), Terraform Associate, ITIL, or related SRE/DevOps certifications.
Licenses, Certifications &
- Training
- N/A
Knowledge, Skills, Abilities, Behaviors
- Strong leadership, consulting, and stakeholder management skills with the ability to influence technical strategy, drive reliability initiatives, and communicate complex technical concepts to engineering and business leaders.
- Exceptional analytical, troubleshooting, and problem-solving abilities with expertise in diagnosing complex production issues, assessing technical risks, and implementing sustainable solutions.
- Demonstrated ability to mentor and guide senior engineers, foster collaboration across teams, and promote Site Reliability Engineering best practices and operational excellence.
- Ability to operate effectively in high-pressure environments, manage competing priorities, and make sound decisions during critical incidents while maintaining a focus on business outcomes.
- Solid ownership mindset, adaptability, and commitment to continuous improvement, with a focus on automation, innovation, reliability, customer experience, and long-term operational maturity.
📌 Senior Consultant - Site Reliability Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad