27 Aug
|
Tavant
|
Bengaluru
- 8+ years of experience in Application Support, Production Support, Application SRE, or Site Reliability Engineering.
- Robust experience supporting Java and/or Node.js applications in production environments.
- Hands-on experience with AWS, Kubernetes, Docker, and Terraform.
- Good understanding of microservices, REST APIs, Linux, networking, and distributed systems.
- Experience with CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI/CD, or similar).
- Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, Datadog, ELK, or Splunk.
- Experience in production deployments, incident management, root cause analysis (RCA), and application troubleshooting.
Production Support &
- Incident Management :
- Monitor the health, availability, and performance of production applications and services.
- Investigate, troubleshoot, and resolve production incidents within agreed SLAs/SLOs.
- Participate in on-call rotations and provide timely support for critical production issues.
- Perform root cause analysis (RCA) and implement preventive and corrective actions to avoid recurring incidents.
- Coordinate with development, infrastructure, and business teams during major incidents.
Application Reliability &
- Operations:
- Ensure high availability, reliability, and stability of Java and Node.js applications running in production.
- Monitor application logs, metrics, and alerts to proactively identify and resolve issues.
- Create and maintain operational runbooks, knowledge articles, and support documentation.
- Continuously improve system observability using monitoring, logging, and alerting tools.
- Support production deployments, configuration changes, and release activities.
Cloud &
- Platform Support:
- Support applications deployed on AWS using containerized environments such as Kubernetes.
- Troubleshoot infrastructure and platform issues related to AWS, Kubernetes, networking, and Terraform-managed resources.
- Work with DevOps teams to automate operational activities and improve deployment processes.
- Ensure application environments are secure, stable, and compliant with organizational standards.
Application Troubleshooting:
- Analyze application issues in Java and Node.js services by reviewing logs, stack traces, API responses, and database interactions.
- Perform basic code analysis and debugging to identify root causes and coordinate fixes with development teams.
- Support API integrations, microservices, messaging systems, and backend services.
Quality &
- Continuous Improvement:
- Identify recurring operational issues and recommend automation or process improvements.
- Develop scripts and automation to reduce manual operational effort and improve system reliability.
- Participate in post-incident reviews and contribute to continuous service improvement initiatives.
Domain &
- Integrations:
- Support multimedia content processing applications, including asset management, moderation, and fraud detection workflows.
- Monitor integrations with external vendors and AI/ML services, ensuring reliable data exchange and service availability.
- Ensure data handling complies with security, privacy, and regulatory requirements.
Collaboration &
- Communication:
- Work closely with Engineering, DevOps, Infrastructure, Product, and Business teams to resolve production issues.
- Communicate incident status, risks, and resolution updates to stakeholders.
- Share operational knowledge and mentor junior support engineers where required.
Disclaimer: This has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.
📌 Backend Engineer - App Support (Bengaluru)
🏢 Tavant
📍 Bengaluru