02 Oct
|
Aziro
|
Bengaluru
Job Summary
Staff Software Engineer, Site Reliability Platform Automation. We have an prospect for a Staff Software Engineer, Site Reliability Platform Automation to join our SaaS Platform Engineering team in Bangalore, India, reporting to the Sr. Manager, Site Reliability Platform Engineering.
You will own the technical delivery of a functional area across global cloud networking and SaaS platforms. You will work hands-on across reliability measurement or toil automation, building systems that reduce operational effort and enable product engineering teams to operate more independently. One opening is anchored in reliability measurement, including SLO and SLI definition across the service catalog, error budget policy and burn alerting, and the observability practice and evidence store that make reliability claims verifiable.
The second is anchored in toil automation, including automated CVE remediation across four production realms, test automation for third-party provider software, and automated EKS and RDS upgrades. You will collaborate with DevOps, CloudOps, product engineering, architecture, security, and product management teams. You will also help apply AI-assisted engineering and operational tools responsibly to improve productivity, analytics, automation, and incident decision support.
Location Bangalore, India
Be Prepared - What You Bring
- 8+ years of software engineering experience, with meaningful depth in infrastructure, platform, reliability, DevOps, or related systems
- Strong production coding ability in Go, Python, or a comparable language, together with infrastructure-as-code fluency, particularly Terraform
- Deep practical Kubernetes experience, including operating, debugging, and improving real production clusters
- Experience with AWS or GCP
- Operational judgment demonstrated through on-call experience, incident participation, production troubleshooting, and sound decision-making under pressure
- A track record of replacing recurring manual work with software and explaining the measured improvement before and after automation
- Experience defining SLIs and SLOs for services owned by other teams and negotiating meaningful targets with service owners
- Nice to have policy-as-code experience with Kyverno, OPA, or Gatekeeper
- Observability stack experience with Prometheus, Grafana, Loki, Cortex, OpenTelemetry, ELK, Datadog, PagerDuty, or comparable technologies
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Staff S/W Engg, SRE & Platform Automation Engg (Bengaluru)
🏢 Aziro
📍 Bengaluru