Lead Site Reliability Engineer/ Expert (Bengaluru)

Lead Site Reliability Engineer/ Expert (Bengaluru)

09 Oct
|
SITA
|
Bengaluru

09 Oct

SITA

Bengaluru

Job Summary

We are seeking a hands-on Lead Site Reliability Engineer (SRE) with solid expertise across application support, Kubernetes environments, and CI/CD pipelines. This role is responsible for ensuring the reliability, performance, and observability of production systems through deep technical analysis and proactive engineering. The successful candidate will lead incident and problem investigations, identify root causes, and drive permanent resolutions in close collaboration with Development, Product, and Operations teams. This is an engineering-focused role requiring strong troubleshooting skills, production support experience, and a commitment to continuous service improvement.

Responsibilities

- Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.
- Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure.
- Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.
- Serve as the technical escalation point during critical incidents, providing clear and timely guidance.
- Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.
- Enhance observability and alerting to improve issue detection and resolution.
- Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.




- Automate repetitive operational tasks and promote engineering best practices.
- Assess the impact of deployments on production environments through CI/CD pipeline expertise.
- Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.

Qualifications

- Min 8 years experience in SRE, DevOps, or Production Engineering supporting high-availability systems.
- Strong expertise in root cause analysis (RCA) and permanent issue resolution.
- Hands-on troubleshooting of production environments using logs, metrics, and traces.
- Strong knowledge of Java and/or .NET applications in production.
- Hands-on experience with Kubernetes and containerized workloads.
- Experience with monitoring and observability tools for distributed systems.
- Familiarity with CI/CD pipelines and deployment tools (e.g., Azure DevOps, Jenkins, GitHub Actions).
- Scripting and automation skills using Python, Bash, or similar.
- Strong analytical, communication, and cross-functional collaboration skills.
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Lead Site Reliability Engineer/ Expert (Bengaluru)
🏢 SITA
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead site reliability engineer/ expert (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: lead site reliability engineer/ expert (bengaluru) / bengaluru