27 Aug
|
Tavant
|
Bengaluru
We are seeking an experienced Application Site Reliability Engineer (Application SRE) to support and maintain mission-critical production applications.
The ideal candidate will have strong expertise in Java and/or Node.js, along with hands-on experience in AWS, Kubernetes, and Terraform.
This role focuses on ensuring application reliability, availability, and operational excellence through proactive monitoring, production support, incident management, automation, and continuous service improvement.
Production Support & Incident Management:
- Monitor, troubleshoot, and support business-critical production applications.
- Investigate and resolve production incidents within agreed SLAs/SLOs.
- Participate in on-call rotations and major incident management.
- Perform Root Cause Analysis (RCA) and implement preventive actions.
Application Reliability & Operations:
- Ensure the availability, stability, and performance of Java and Node.js applications.
- Monitor application health using logs, metrics, dashboards, and alerts.
- Create and maintain operational runbooks and support documentation.
- Support production deployments, configuration changes, and release activities.
Cloud & Platform Support:
- Support applications running on AWS and Kubernetes platforms.
- Troubleshoot cloud infrastructure and container-related issues.
- Manage infrastructure using Terraform and support CI/CD deployment pipelines.
- Improve application observability and operational automation.
Collaboration & Continuous Improvement:
- Work closely with Development, DevOps, Platform Engineering, and Product teams to resolve production issues.
- Identify recurring issues and drive automation to reduce operational effort.
- Contribute to operational excellence by improving monitoring, alerting, and support processes.
Qualifications & Requirements:
- 8+ years of experience in Application Support, Production Support, Application SRE, or Site Reliability Engineering.
- Strong experience supporting Java and/or Node.js applications in production environments.
- Hands-on experience with AWS, Kubernetes, Docker, and Terraform.
- Robust understanding of Linux, Microservices, REST APIs, and distributed systems.
- Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, Datadog, ELK, or Splunk.
- Experience in production deployments, incident management, troubleshooting, and Root Cause Analysis (RCA).
- Familiarity with multimedia formats and image/video processing pipelines is preferred.
- Excellent written and verbal communication skills
Disclaimer: This has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.
📌 Backend Site Reliability Engineer (Bengaluru)
🏢 Tavant
📍 Bengaluru