Sre production support and Application Support (Hyderabad)

Sre production support and Application Support (Hyderabad)

09 Oct
|
VY SYSTEMS PRIVATE
|
Hyderabad

09 Oct

VY SYSTEMS PRIVATE

Hyderabad

ob Summary

We are looking for an experienced Site Reliability Engineer (SRE) / Production Support Engineer with robust hands-on experience in application and production support, incident management, monitoring, cloud operations, automation, and infrastructure technologies.

The ideal candidate will be responsible for ensuring the availability, reliability, performance, and stability of production applications and infrastructure. The role involves troubleshooting critical production issues, monitoring applications and infrastructure, supporting deployments, managing incidents, and driving automation and operational improvements.

Key Responsibilities

- Provide L2/L3 Application and Production Support for critical business applications.
- Monitor production applications, infrastructure, batch jobs, and system health.
- Handle and troubleshoot critical production incidents, ensuring timely resolution and minimal business impact.
- Participate in Incident, Problem, and Change Management processes.
- Perform root-cause analysis (RCA) for recurring and major production issues.
- Troubleshoot issues related to applications, networks, load balancers, databases, operating systems, and infrastructure.
- Monitor application and infrastructure performance using Splunk, APM, and other monitoring tools.
- Create and maintain Splunk queries, dashboards, alerts, and operational monitoring.
- Support production deployments, including Blue-Green and Canary deployment strategies.
- Work with cloud infrastructure and perform day-to-day Cloud Operations activities.




- Manage and troubleshoot containerized applications using Docker and Kubernetes.
- Work with Terraform / Infrastructure as Code (IaC) for infrastructure provisioning and automation.
- Support Linux and Windows server administration.
- Develop and maintain Shell scripts and Python automation scripts to reduce manual operational activities.
- Monitor and analyze SLIs, SLOs, Error Budgets, and Burn Rates.
- Identify reliability risks and proactively implement solutions to improve system availability and performance.
- Collaborate with Development, DevOps, Infrastructure, Network, Database, and Cloud teams during production incidents.
- Leverage GenAI tools such as GitHub Copilot, Claude, or similar tools to improve troubleshooting, automation, documentation, and operational efficiency.
- Participate in on-call/shift-based production support as required by business and customer needs.

Mandatory / Key Skills

- Production / Application Support
- Incident Management
- Production Monitoring & Batch Monitoring
- Splunk – Queries, Dashboards & Monitoring
- APM / Application Performance Monitoring
- SLI / SLO / Error Budget / Burn Rate
- Cloud Operations
- Kubernetes
- Docker
- Terraform / Infrastructure as Code
- Linux & Windows Administration
- Shell Scripting
- Python Scripting / Automation
- Production Deployment Support
- Blue-Green & Canary Deployments
- Network, Load Balancing & Database Troubleshooting

​

Skills:- Production support, Application server, Kubernetes, Terraform, Amazon Web Services (AWS), Python and Splunk

📌 Sre production support and Application Support (Hyderabad)
🏢 VY SYSTEMS PRIVATE
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sre production support and application support (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: sre production support and application support (hyderabad) / hyderabad