24 Sep
|
Aziro
|
Bengaluru
Senior Staff Software Engineer, Site Reliability & Platform Automation
We have an prospect for a Senior Staff Software Engineer, Site Reliability & Platform Automation to join our SaaS Platform Engineering team in Bangalore, India, reporting to Sr. Manager, Site Reliability & Platform Engineering. In this pivotal role, you will set technical direction across reliability measurement, toil automation, resilience testing, and self-service enablement for Infobloxs global cloud networking and SaaS platforms.
Collaborating closely with DevOps, CloudOps, product engineering, architecture, security, and product management teams, you will design and build platforms that improve reliability, reduce operational toil, and help engineering teams own more of their DevOps lifecycle. You will also champion the responsible use of AI-assisted tools for engineering productivity, operational analytics, automation, content generation, and incident decision support.
Be a Contributor — What You’ll Do
- Set technical direction across reliability measurement, toil automation, resilience testing, and self-service enablement, ensuring the architecture operates as a coherent platform rather than four disconnected efforts
- Carry the technical design across High Availability, SLO/SLI, Error Budgets, Chaos/DR Testing, Change Management Governance, Incident Response, Observability, and Application Team Performance & Capacity Tuning
- Design and build golden paths, Terraform modules, and Kyverno policies that enable product engineering teams to provision and operate services independently through guardrails rather than approval gates
- Design measurement pipelines and evidence stores that make reliability, service health, developer experience, operational efficiency, and platform adoption measurable and auditable
- Build automation frameworks for repeatable CVE remediation, infrastructure upgrades, third-party provider testing, and routine operational requests across four production realms
- Partner with NA DevOps, IN DevOps,
and CloudOps Engineering teams to understand manual workflows and replace recurring operational work with reliable, maintainable software
- Evaluate build-versus-adopt decisions and select appropriate open-source, commercial, or internal capabilities to avoid creating unnecessary platform toil
- Develop and review production-quality software in Go, Python, Rust, Java, or a comparable language, including tests, instrumentation, documentation, and operational safeguards
- Establish resilient, secure, and observable designs for AWS and GCP environments, including platforms with FedRAMP, SOC 2, or ISO-related constraints
- Apply AI-assisted engineering tools to accelerate code exploration, test generation, documentation, telemetry analysis, workflow automation, and operational decision support while preserving human review, security, privacy, auditability, and accountability
Be Prepared — What You Bring
- 12+ years of software engineering experience, with substantial depth in infrastructure, platform, developer tooling, distributed systems, or reliability engineering
- Demonstrated experience building and evolving internal platforms, developer tools, or reliability systems used by other engineering teams in production
- Deep expertise in at least two of distributed systems, Kubernetes and container platforms, infrastructure as code, CI/CD systems, observability and telemetry pipelines, with working fluency across the remaining areas
- Strong current production coding ability in Go, Python, Rust, Java, or a comparable language, including testing, debugging, performance analysis, and operational ownership
- Experience designing and operating cloud platforms across AWS and/or GCP, with practical Kubernetes experience and infrastructure as code proficiency, particularly Terraform
- Experience establishing or scaling SLO/SLI, error budget, observability, incident response, high availability, disaster recovery, or resilience testing practices
- Proven ability to replace recurring manual work with automation and demonstrate measurable improvements in toil, reliability, capacity, change safety, or developer experience
- Platform-as-product mindset, with experience treating internal engineers as customers and measuring adoption, usability, satisfaction, and operational outcomes
- Practical experience using AI-assisted engineering or operational tools for productivity, analytics, automation, content generation, or decision support, together with sound judgment about appropriate human oversight and control boundaries
- Proven ability to influence senior technical stakeholders, mentor engineers, communicate clearly, and drive adoption across organizational boundaries; a bachelor’s or master’s degree in computer science, engineering, or a related technical field is preferred
Nice to have
- Experience operating multi-tenant, multi-region, or multi-cloud Kubernetes and cloud platforms at organizational scale
- Experience working with regulated environments or controls such as FedRAMP, SOC 2, or ISO
- Experience with Prometheus, Grafana, OpenTelemetry, Loki, ELK, Datadog, or comparable observability platforms
- Experience with policy as code, including Kyverno, OPA, or Gatekeeper
- Experience with Jenkins, GitHub Actions, GitOps, Argo, or comparable CI/CD technologies
- Experience with chaos engineering, fault injection, disaster recovery game days, regional failover, or resilience testing at production scale
- Experience applying LLM-based or agentic tooling to operational workflows with appropriate safeguards
📌 Senior Staff Software Engineer, Site Reliability & Platform Automation (Bengaluru)
🏢 Aziro
📍 Bengaluru