07 Aug
|
Softenger
|
India
Job Responsibilities:
Design, implement, and maintain highly reliable, scalable, and observable systems on Microsoft Azure. Define and operationalize SLIs, SLOs, and SLAs to measure and improve service reliability. Build and enhance observability platforms using tools like Grafana, Prometheus, ELK stack, and OpenTelemetry. Drive adoption of SRE principles (error budgets, toil reduction, automation-first mindset). Implement proactive monitoring, alerting, and incident response frameworks. Lead incident management, root cause analysis (RCA), and postmortems with a blameless culture. Automate infrastructure and workflows using Infrastructure as Code (IaC). Collaborate with engineering teams to improve system resilience, performance, and deployment practices. Develop and maintain runbooks, playbooks, and operational standards. Advocate for DevOps and SRE culture adoption across teams
Desired Skill:
SRE Practices Hands-on experience implementing: o SLIs, SLOs,
SLAs o Error budgets o Toil reduction strategies Strong understanding of incident management lifecycle
Relevant Exp:
Programming & Automation Proficiency in Python (automation, tooling, scripting) Experience building internal tools for reliability and observability Infrastructure as Code (IaC) Strong experience with Terraform Familiarity with infrastructure automation and configuration management Experience with GitHub Actions (or similar CI/CD tools) Knowledge of deployment strategies (blue-green, canary, rolling updates)
Value Add:
Valuable to Have Experience with AI/ML Observability (monitoring models, drift detection, LLM observability) Familiarity with: o Service Mesh (Istio, Linkerd) o Chaos Engineering tools (e.g., Chaos Monkey, Litmus) o Distributed tracing tools (Jaeger, Tempo) Exposure to FinOps practices (cost optimization in cloud) Experience with multi-cloud or hybrid environments
📌 Site Reliability Engineer (India)
🏢 Softenger
📍 India