15 Sep
|
HCA Healthcare UK
|
Hyderabad
15 Sep
HCA Healthcare UK
Hyderabad
Senior Consultant - Site Reliability Engineering (SRE)
Experience
11+ years
About This Role The Senior Consultant - Site ReliabilityEngineering (SRE) is a hands-on technical leadership and consulting roleresponsible for driving the reliability, availability, performance, resilience,and operational maturity of enterprise and business-critical services. The roleserves as a principal technical escalation point and trusted advisor forcomplex production challenges, shaping SRE strategy and engineeringimprovements across observability, automation, Infrastructure-as-Code, incidentmanagement, reliability engineering, performance optimization, and operationalreadiness.
The Senior
Consultant partners with Development, Architecture, SRE,DevOps, Cloud, Database, Network, Infrastructure, Security, and businessstakeholders; leads cross-functional initiatives; and mentors senior engineers.
Essential Duties
Serve as a principal technical escalation pointand SRE consultant for complex, high-impact, and business-critical incidents;lead cross-functional troubleshooting, executive-level technical communication,and service restoration.
Troubleshoot across applications, database,API/integration, middleware, cloud, server, network, identity, security, andexternal dependency layers using logs, metrics, traces, events, andinfrastructure telemetry.
Lead root cause analysis for significantincidents and drive corrective and preventive actions that reduce recurrenceand operational risk.
Define, govern, and mature SRE and observabilitypractices, including SLIs/SLOs, error budgets, dashboards, alerting,instrumentation, event correlation, monitoring coverage, and reliabilityreporting.
Lead enterprise reliability , operational-readiness , and supportability assessments; identify gaps inresiliency, automation, observability, documentation, infrastructure, anddeployment processes and develop prioritized multi-quarter improvement roadmaps.
Lead strategic toil-reduction programs anddesign reusable automation, self-healing, and auto-remediation patterns thatimprove engineering productivity and service reliability at scale.
Provide technical governance and hands-onleadership for Infrastructure-as-Code, configuration management, GitOps,platform engineering, and CI/CD practices using technologies such as Terraform,Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
Troubleshoot complex deployment, configuration,pipeline, rollback, and release failures and partner with engineering teams toimprove deployment reliability.
Support major application upgrades, migrations,platform modernization, patching, environment transitions, and productioncutovers.
Analyze application and infrastructureperformance / capacity trends and recommend scaling, quota, configuration,resiliency, and cost-optimization improvements.
Own technical direction for disaster recovery,business continuity engineering, resiliency validation,
recoveryobjectives/procedures, failure testing, and rollback readiness.
Partner with Security and engineering teamsduring critical vulnerabilities or cyber events and support application,infrastructure, authentication, and configuration analysis.
EstablishSRE standards, reference architectures, playbooks, runbooks, templates,operating procedures, governance mechanisms, and reusable engineering patternsacross teams.
Evaluateemerging technologies and operating practices in observability, automation,cloud operations, reliability engineering, platform engineering, andAI-assisted operations; lead proof-of-concept evaluations, technicalrecommendations, and adoption roadmaps.
Mentorsenior SRE, Production Engineering, DevOps, and Application Support engineers;provide technical coaching in troubleshooting, observability, automation,incident management, root cause analysis, architecture, and reliabilitypractices.
Useincident trends, operational metrics, SLO performance, risk indicators, andengineering data to define reliability priorities, influence stakeholders, anddrive measurable continuous improvement.
Provide consultative leadership to applicationand platform teams on reliability architecture, SRE adoption, productionreadiness, cloud modernization, and operational risk reduction.
Lead reliability reviews with seniorstakeholders, translate technical risk into business impact, and definemeasurable remediation plans, success criteria, and governance checkpoints.
Participate inan on-call or senior production escalation rotation where required.
PositionRequirements
11+years of progressive experience in Site Reliability Engineering, ProductionEngineering, Application/Production Support, DevOps, Cloud/PlatformEngineering, or a related technical discipline, including significantexperience leading complex enterprise reliability initiatives.
Strong hands-on experience supporting enterpriseor business-critical production applications and troubleshooting acrossmultiple technology layers.
Advanced experience with monitoring, logging,APM, and observability platforms such as Dynatrace, Splunk, Grafana, orcomparable technologies.
Solid working knowledge of Windows and/orLinux/Unix platforms, relational databases, SQL, APIs, integrations, anddistributed application dependencies.
Experience with cloud platforms such as Azure,GCP, AWS, or comparable enterprise cloud technologies.
Hands-on experience with CI/CD, source control,GitOps/deployment tools, Infrastructure-as-Code,
and automation technologiessuch as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible.
Strong understanding of networking conceptsincluding DNS, firewall rules, ports, load balancing, routing, certificates,and application connectivity.
Strong understanding of identity, privilegedaccess, service accounts, secrets, and enterprise security concepts.
Demonstrated experience leading complex incidenttroubleshooting, root cause analysis, and implementation ofcorrective/preventive improvements.
Demonstrated ability to act as a seniortechnical consultant, mentor experienced engineers, influence architecture andengineering decisions, communicate effectively with technical and executivestakeholders, and drive cross-functional improvements without formalpeople-management authority.
Education
Bachelorsdegree in Computer Science, Information Technology, Engineering, or a relateddiscipline preferred; equivalent advanced technical experience and demonstratedSRE leadership may be considered.
Relevant certifications in Cloud, DevOps, SRE,Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud,or related areas are beneficial.
Knowledge andSkills
Capability
Knowledge & Skill Expectation
Application & Production Engineering
Advanced troubleshooting across complex applications, integrations, dependencies, and production environments.
Observability & Reliability
Expert knowledge of SRE principles, SLIs/SLOs, error budgets, logs, metrics, traces, dashboards, alerting, performance, availability, resilience, and reliability governance.
Cloud & Infrastructure
Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies.
DevOps & Infrastructure-as-Code
Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices.
Automation & Engineering
Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows.
Incident & Problem Management
Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements.
Performance, Capacity & Security
Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams.
Technical Leadership & Improvement
Expert ability to operate as a senior SRE consultant, mentor experienced engineers, influence architecture and stakeholders, establish enterprise standards, lead transformation roadmaps, and drive measurable reliability improvements.
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Consultant - Site Reliability Engineer (Hyderabad)
🏢 HCA Healthcare UK
📍 Hyderabad