Senior - System Engineer (Hyderabad)

Senior - System Engineer (Hyderabad)

07 Oct
|
HCA Healthcare - India
|
Hyderabad

07 Oct

HCA Healthcare - India

Hyderabad

General Position Information

Reports directly to (Title): Senior – Platform Engineer

Matrix reports to (Title): As applicable based on functional alignment

Direct Reports: Individual contributor; no direct reports

Created / Last Revised: -

Job Code: -

Position Summary The Application Support Engineer is a hands-on technical role responsible for the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical applications. The role serves as a senior technical escalation point for complex production issues and drives improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The engineer partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, and Security teams and mentors other support engineers.

Responsibilities

- Serve as a technical escalation point for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting through service restoration.
- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
- Define and improve application reliability and observability practices, including service indicators/objectives, dashboards, alerting, instrumentation, event correlation, and monitoring coverage.
- Conduct application reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop improvement roadmaps.
- Lead toil-reduction efforts and design reusable automation, self-healing, and auto-remediation workflows for repetitive operational activities.
- Guide and support Infrastructure-as-Code, configuration management, GitOps, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
- Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers.




- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.
- Provide technical leadership for disaster recovery, resiliency validation, recovery procedures, and rollback readiness.
- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
- Establish and improve technical standards, playbooks, runbooks, templates, operating procedures, and reusable support patterns.
- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, and AI-assisted operations; contribute to proof-of-concept evaluations and recommendations.
- Mentor Application Support Engineers in troubleshooting, monitoring, automation, incident management, root cause analysis, and reliability practices.
- Use incident trends, support metrics, reliability indicators, and operational data to drive measurable continuous improvement.
- Participate in an on-call or senior production escalation rotation where required.

Education & Experience

- Bachelor’s degree in computer science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience may be considered.
- 4+ years of experience in Application Support, Production Engineering, Site Reliability Engineering, DevOps, Cloud Operations, or a related technical discipline.

Must Have Skills

- Strong hands-on experience supporting enterprise or business-critical production applications and troubleshooting across multiple technology layers.
- Advanced experience with monitoring, logging, APM, and observability platforms such as Dynatrace, Splunk, Grafana, or comparable technologies.
- Strong working knowledge of Windows and/or Linux/Unix platforms, relational databases, SQL, APIs, integrations, and distributed application dependencies.
- Experience with cloud platforms such as Azure, GCP, AWS, or comparable enterprise cloud technologies.
- Hands-on experience with CI/CD, source control, GitOps/deployment tools, Infrastructure-as-Code,



and automation technologies such as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible.
- Strong understanding of networking concepts including DNS, firewall rules, ports, load balancing, routing, certificates, and application connectivity.
- Strong understanding of identity, privileged access, service accounts, secrets, and enterprise security concepts.
- Demonstrated experience leading complex incident troubleshooting, root cause analysis, and implementation of corrective/preventive improvements.
- Demonstrated ability to mentor engineers, influence cross-functional teams, communicate technical findings, and drive engineering improvements without formal people-management authority.

Nice To Have Skills

- Experience with Kubernetes, container orchestration platforms, and cloud-native application support.
- Hands-on experience with scripting or programming languages such as PowerShell, Python, Bash, C#, or Java for automation and operational tooling.
- Knowledge of SRE principles including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability engineering practices.
- Experience supporting microservices architectures, event-driven systems, and enterprise integration platforms.
- Familiarity with AIOps, predictive monitoring, automated remediation, and contemporary observability practices using machine learning-assisted operations tools.

Licenses, Certifications & Training

- Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.

Knowledge, Skills, Abilities, Behaviors

- Strong analytical, troubleshooting, and problem-solving skills with the ability to diagnose and resolve complex technical issues across applications, infrastructure, cloud platforms, databases, and integrations.
- Excellent communication, collaboration, and stakeholder management skills with the ability to work effectively across engineering, operations, security, and business teams.
- Demonstrated ability to lead incident response, drive root cause analysis, and implement sustainable corrective and preventive actions that improve service reliability.
- Strong commitment to operational excellence, automation, continuous improvement, and adoption of Site Reliability Engineering (SRE) best practices.
- Ability to manage multiple priorities in a fast-paced environment while demonstrating ownership, accountability, adaptability, and a customer-focused mindset.

📌 Senior - System Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior - system engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: senior - system engineer (hyderabad) / hyderabad