26 Aug
|
NTT DATA Business Solutions
|
Bengaluru
26 Aug
NTT DATA Business Solutions
Bengaluru
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) with a strong background in Kubernetes, production support, observability, SQL, and Java-based applications. The role will focus on ensuring the availability, reliability, scalability, and performance of business-critical production systems.
We are currently seeking a SRE Reliability Engineer to join our team in Bangalore, Karnataka, India.
The ideal candidate will combine robust application troubleshooting skills with SRE and DevOps practices, using Datadog and/or Prometheus for monitoring and observability and Kubernetes for managing containerized workloads. A strong understanding of Java applications and relational databases/SQL is essential for diagnosing issues across application, infrastructure, and data layers.
Key Responsibilities
- Own the reliability, availability, and operational health of business-critical production applications and services.
- Provide L2/L3 production support, including incident triage, troubleshooting, resolution, and stakeholder communication.
- Monitor and support applications deployed on Kubernetes, including pods, deployments, services, ingress, resource utilization, scaling, and cluster-related issues.
- Implement and maintain application and infrastructure monitoring using Datadog, Prometheus, dashboards, metrics, logs, and alerts.
- Define and monitor SLIs, SLOs, SLAs, error budgets, and service-health indicators for critical applications.
- Troubleshoot production issues across Java applications, APIs, microservices, Kubernetes, databases, and infrastructure.
- Analyze Java application logs, exceptions, JVM performance, memory utilization, thread behavior, and garbage collection to identify performance and reliability issues.
- Use SQL to investigate production incidents, validate data, identify data-related issues, and perform application-level troubleshooting.
- Participate in incident management, major incident calls, root-cause analysis (RCA), and post-incident reviews.
- Identify recurring production issues and drive permanent remediation through automation and engineering improvements.
- Build and enhance monitoring dashboards, alerting mechanisms, and operational runbooks to improve early detection and reduce recovery time.
- Drive improvements in MTTR, availability, performance, capacity, and production stability.
- Automate repetitive operational activities using scripting and appropriate DevOps/SRE tooling.
- Support application releases, production deployments, rollback activities, and post-deployment validation.
- Work closely with Development, Infrastructure, DevOps, Database, Security, and Business teams to ensure production readiness.
- Participate in on-call and production support rotations as required.
Required Skills & Experience
- 5+ years of overall IT experience, with significant experience in SRE, Production Support, Application Support, or DevOps roles.
- Strong hands-on experience with Kubernetes and containerized applications.
- Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
- Strong experience supporting Java/J2EE or Java-based microservices applications in production.
- Good understanding of JVM troubleshooting, application logs, memory, threads, garbage collection, and performance issues.
- Strong SQL skills with experience troubleshooting relational databases and application data issues.
- Experience managing P1/P2 production incidents, including incident coordination, RCA, and problem management.
- Good understanding of REST APIs, microservices, distributed systems, and application integration patterns.
- Experience with Linux/Unix environments and shell scripting.
- Understanding of CI/CD pipelines, release management, and deployment practices.
- Strong analytical and troubleshooting skills with the ability to diagnose issues across multiple technology layers.
Preferred Skills
- Experience with cloud platforms such as AWS/ Azure.
- Familiarity with Docker, Helm, GitLab/GitHub Actions, or similar DevOps tooling.
- Experience with centralized logging platforms such as ELK/OpenSearch or Splunk.
- Exposure to Infrastructure as Code tools such as Ansible/Terraform.
- Understanding of load balancing, networking, DNS, certificates, and application security.
- Experience implementing automation to reduce manual operational effort and production toil.
- Familiarity with ITIL processes including Incident, Problem, and Change Management.
Key SRE Competencies
- Production Reliability: Ability to maintain highly available and resilient production services.
- Observability: Strong understanding of metrics, logs, traces, dashboards, and actionable alerting.
- Incident Management: Ability to rapidly diagnose and restore services during critical incidents.
- Problem Management: Strong RCA skills with focus on permanent remediation rather than repeated tactical fixes.
- Automation: Ability to identify and automate repetitive production-support activities.
- Performance Engineering: Ability to identify bottlenecks across Java applications, Kubernetes, and databases.
- Stakeholder Management: Ability to communicate clearly during incidents and work effectively across engineering and business teams.
📌 SRE Reliability Engineer (Bengaluru)
🏢 NTT DATA Business Solutions
📍 Bengaluru