Role: Production Support Engineer
nLocation: Hyderabad
nExperience 7-12 years
nApply here:
nAbout the role
nWe’re looking for an Operations / Production Engineer to keep business-critical services stable, secure and available day to day. You’ll combine strong Linux/Unix fundamentals, practical SQL skills and structured incident management to diagnose issues, restore services quickly and improve long-term reliability.
nThis role suits someone who enjoys solving real-world production problems, working calmly under pressure and turning recurring incidents into automation, monitoring and preventative improvements.
nKey responsibilities
nProduction support and reliability
n
n
- Monitor the health, performance and availability of production services.
n
- Investigate and resolve incidents, outages, performance degradation and service alerts.
n
- Perform structured triage, identify business impact and prioritise response accordingly.
n
- Use Linux/Unix tools to diagnose system, process, network, memory and storage issues.
n
- Analyse application and system logs to identify root causes and contributing factors.
n
- Execute approved operational procedures, recovery activities and service restoration plans.
n
- Participate in on-call or out-of-hours support arrangements, where required.
n
nMonitoring and automation
n
n
- Develop and improve monitoring, alerting and service-health checks.
n
- Reduce manual effort through scripting, automation and repeatable operational tooling.
n
- Identify recurring incidents and deliver preventative improvements.
n
- Contribute to capacity, resilience, disaster-recovery and operational-readiness activities.
n
- Improve runbooks, standard operating procedures and knowledge articles.
n
nIncident and problem management
n
n
- Manage incidents from initial report through diagnosis,
escalation, resolution and closure.
n
- Communicate clearly with stakeholders throughout the incident lifecycle.
n
- Escalate to specialist teams and suppliers when required, providing useful evidence and impact details.
n
- Support root-cause analysis and post-incident reviews.
n
- Track corrective and preventative actions through to completion.
n
- Maintain accurate incident, change and problem records.
n
nDatabase and data investigation
n
n
- Excellent knowledge of SQL queries, joins, aggregation etc…
n
- Able to identify non-performing SQL and optmise it in coordination with development team
n
nEssential skills and experience
n
n
- Experience in Operations Engineering, Production Support, Site Reliability Engineering, Infrastructure Support or a similar role.
n
- Understanding of incident, change and problem-management practices aligned to IT service-management principles.
n
- Ability to assess impact, prioritise incidents and work effectively under pressure.
n
- Solid written and verbal communication skills.
n
- A disciplined approach to documentation, risk management and operational controls.
n
- Commitment to security, resilience, service quality and continuous improvement.
n
- Working experience on PostgreSQL and Linux platform is a must
n
nDesirable skills
n
n
- Knowledge of ITIL practices and service-management tooling.
n
- Familiarity with observability platforms, metrics, dashboards, alerting and distributed tracing.
n
- Scripting or automation experience using languages such as Bash, Python or PowerShell.
n
- Experience with deployment pipelines, version control and infrastructure-as-code.
n
- Understanding of resilience testing, disaster recovery and capacity management.
n
- Experience working in regulated, financial-services or other highly controlled environments.
n
n
📌 Production Support Specialist (Sangli)
🏢 CAPCO
📍 Sangli