07 Oct
|
HCA Healthcare - India
|
Hyderabad
07 Oct
HCA Healthcare - India
Hyderabad
General Position Information
Reports directly to (Title): Senior – Platform Engineer
Matrix reports to (Title): As applicable based on functional alignment
Direct Reports: Individual contributor; no direct reports
Created / Last Revised: -
Job Code: -
Position Summary The Application Support Engineer is a hands-on technical role responsible for the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical applications. The role serves as a senior technical escalation point for complex production issues and drives improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The engineer partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, and Security teams and mentors other support engineers.
Responsibilities
- Serve as a technical escalation point for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting through service restoration.
- Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
- Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
- Define and improve application reliability and observability practices, including service indicators/objectives, dashboards, alerting, instrumentation, event correlation, and monitoring coverage.
- Conduct application reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop improvement roadmaps.
- Lead toil-reduction efforts and design reusable automation, self-healing, and auto-remediation workflows for repetitive operational activities.
- Guide and support Infrastructure-as-Code, configuration management, GitOps, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
- Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
- Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers.
- Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.
- Provide technical leadership for disaster recovery, resiliency validation, recovery procedures, and rollback readiness.
- Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
- Establish and improve technical standards, playbooks, runbooks, templates, operating procedures, and reusable support patterns.
- Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, and AI-assisted operations; contribute to proof-of-concept evaluations and recommendations.
- Mentor Application Support Engineers in troubleshooting, monitoring, automation, incident management, root cause analysis, and reliability practices.
- Use incident trends, support metrics, reliability indicators, and operational data to drive measurable continuous improvement.
- Participate in an on-call or senior production escalation rotation where required.
Education & Experience
- Bachelor’s degree in computer science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience may be considered.
- 4+ years of experience in Application Support, Production Engineering, Site Reliability Engineering, DevOps, Cloud Operations, or a related technical discipline.
Must Have Skills
- Strong hands-on experience supporting enterprise or business-critical production applications and troubleshooting across multiple technology layers.
- Advanced experience with monitoring, logging, APM, and observability platforms such as Dynatrace, Splunk, Grafana, or comparable technologies.
- Strong working knowledge of Windows and/or Linux/Unix platforms, relational databases, SQL, APIs, integrations, and distributed application dependencies.
- Experience with cloud platforms such as Azure, GCP, AWS, or comparable enterprise cloud technologies.
- Hands-on experience with CI/CD, source control, GitOps/deployment tools, Infrastructure-as-Code,
and automation technologies such as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible.
- Strong understanding of networking concepts including DNS, firewall rules, ports, load balancing, routing, certificates, and application connectivity.
- Strong understanding of identity, privileged access, service accounts, secrets, and enterprise security concepts.
- Demonstrated experience leading complex incident troubleshooting, root cause analysis, and implementation of corrective/preventive improvements.
- Demonstrated ability to mentor engineers, influence cross-functional teams, communicate technical findings, and drive engineering improvements without formal people-management authority.
Nice To Have Skills
- Experience with Kubernetes, container orchestration platforms, and cloud-native application support.
- Hands-on experience with scripting or programming languages such as PowerShell, Python, Bash, C#, or Java for automation and operational tooling.
- Knowledge of SRE principles including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability engineering practices.
- Experience supporting microservices architectures, event-driven systems, and enterprise integration platforms.
- Familiarity with AIOps, predictive monitoring, automated remediation, and contemporary observability practices using machine learning-assisted operations tools.
Licenses, Certifications & Training
- Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.
Knowledge, Skills, Abilities, Behaviors
- Strong analytical, troubleshooting, and problem-solving skills with the ability to diagnose and resolve complex technical issues across applications, infrastructure, cloud platforms, databases, and integrations.
- Excellent communication, collaboration, and stakeholder management skills with the ability to work effectively across engineering, operations, security, and business teams.
- Demonstrated ability to lead incident response, drive root cause analysis, and implement sustainable corrective and preventive actions that improve service reliability.
- Strong commitment to operational excellence, automation, continuous improvement, and adoption of Site Reliability Engineering (SRE) best practices.
- Ability to manage multiple priorities in a fast-paced environment while demonstrating ownership, accountability, adaptability, and a customer-focused mindset.
📌 Senior - System Engineer (Hyderabad)
🏢 HCA Healthcare - India
📍 Hyderabad