01 Sep
|
Rabi Solutions
|
Bengaluru
01 Sep
Rabi Solutions
Bengaluru
We are looking for an experienced Senior Site Reliability Engineer (SRE) to join our platform reliability team. The ideal candidate should have robust hands-on experience in Kubernetes, observability, production operations, incident management, and canary deployments.
? Key Responsibilities:
- Maintain Dynatrace Davis AI alerting and service-level problem notification routing
- Configure custom metric event alerting and SLO/burn-rate monitoring
- Build and maintain production platform runbooks
- Handle Vault sealed recovery, Kubernetes node failures, and Kafka broker failures
- Participate in the on-call rotation and lead incident response activities
- Handle P1 incident escalation, post-mortems, 5-Why analysis and corrective actions
- Operate Argo Rollouts and monitor canary deployments
- Perform manual rollback and recovery when automated mechanisms do not trigger
- Monitor AnalysisRun results and overall canary health
? Key Skills:
Kubernetes | SRE | Dynatrace | Davis AI | Argo Rollouts | Kafka | HashiCorp Vault | Observability | SLO | Incident Management | On-Call | Canary Deployments
? Ideal Profile:
- 10+ years of experience in SRE / DevOps / Platform Engineering, with strong experience managing highly available production environments.
📌 Senior Site Reliability Engineer (SRE) – Kubernetes & Observability (Bengaluru)
🏢 Rabi Solutions
📍 Bengaluru