24 Sep
|
Aziro
|
Bengaluru
Staff Software Engineer, Site Reliability & Platform Automation:
We have an opportunity for a Staff Software Engineer, Site Reliability & Platform Automation to join our SaaS Platform Engineering team in Bangalore, India, reporting to the Sr. Manager, Site Reliability & Platform Engineering. You will own the technical delivery of a functional area across our global cloud networking and SaaS platforms.
You will work hands-on across reliability measurement or toil automation, building systems that reduce operational effort and enable product engineering teams to operate more independently.
One opening is anchored in reliability measurement, including SLO and SLI definition across the service catalog, error budget policy and burn alerting, and the observability practice and evidence store that make reliability claims verifiable. The second is anchored in toil automation, including automated CVE remediation across four production realms, test automation for third-party provider software, and automated EKS and RDS upgrades.
You will collaborate with DevOps, CloudOps, product engineering, architecture, security, and product management teams. You will also help apply AI-assisted engineering and operational tools responsibly to improve productivity, analytics, automation, and incident decision support.
Be a Contributor What Youll Do
- Own a functional area end to end, including architecture, implementation, rollout, adoption, operational health, and evolution
- Design and build production software that automates infrastructure and reliability workflows; spend most of your time writing and reviewing software
- For the reliability measurement opening, define and implement SLIs, SLOs, error budgets, burn-rate alerting, observability standards, and a durable evidence store across critical and core services
- For the toil automation opening,
build automated workflows for CVE remediation across four realms, EKS and RDS upgrades, third-party provider testing, and recurring infrastructure operations
- Work directly with NA DevOps, IN DevOps, and CloudOps engineers to understand manual workflows before replacing them with reliable, measurable automation
- Build self-service capabilities wherever product engineering teams can safely operate them independently, using guardrails rather than approval gates
- Establish measurement baselines for toil, reliability, adoption, developer experience, and operational cost before changing systems or workflows
- Build golden paths, Terraform modules, policy-as-code controls, and reusable platform components that shift DevOps ownership toward the teams that build and own the services
- Partner with product engineering teams to improve production readiness, observability, incident response, capacity planning, change safety, and operational ownership
- Land adoption across peer engineering and product engineering teams; treat usage, satisfaction, and reduced manual effort as part of delivery
Be Prepared What You Bring
- 8+ years of software engineering experience, with meaningful depth in infrastructure, platform, reliability, DevOps, or related systems
- Experience owning a significant production system, including its failure modes, operational cost, reliability, and evolution
- Solid production coding ability in Go, Python, or a comparable language, together with infrastructure-as-code fluency, particularly Terraform
- Deep practical Kubernetes experience, including operating, debugging, and improving real production clusters
- Experience with AWS or GCP, CI/CD systems, distributed systems, and production operations
- Operational judgment demonstrated through on-call experience, incident participation, production troubleshooting, and sound decision-making under pressure
- A track record of replacing recurring manual work with software and explaining the measured improvement before and after automation
- Experience building internal platforms, developer tooling, or reliability capabilities that other engineering teams adopted
- For the reliability measurement opening: experience defining SLIs and SLOs for services owned by other teams and negotiating meaningful targets with service owners
- For the toil automation opening: experience with vulnerability remediation at scale, automated infrastructure upgrades, or test automation for third-party software that you cannot modify
Nice to have
- Policy-as-code experience with Kyverno, OPA, or Gatekeeper
- Observability stack experience with Prometheus, Grafana, Loki, Cortex, OpenTelemetry, ELK, Datadog, PagerDuty, or comparable technologies
- CI/CD platform engineering experience, including Jenkins, GitHub Actions, GitOps, Argo, or large-scale CI/CD platform migration
- Experience with multi-tenant, multi-region, or regulated environments, including FedRAMP, SOC 2, or ISO-controlled platforms
- Internal developer platform or developer experience work with measurable adoption outcomes
- Experience applying LLM-based or agentic tools to engineering or operational workflows, including human approval and controlled remediation
- Experience with EKS, RDS, PostgreSQL, cloud networking, or disaster recovery testing
Location Bangalore, India
📌 Staff Software Engineer, Site Reliability & Platform Automation (Bengaluru)
🏢 Aziro
📍 Bengaluru