A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.
Key Responsibilities
- Own the reliability of production systems for one or more enterprise customers on AWS.
- Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
- Lead the response to reliable (P0/P1) incidents and run blameless post-mortems that lead to fixes.
- Lead within the on-call rotation.
- Operate and optimise Kubernetes clusters and AWS services as workloads grow.
- Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
- Advise customers on improving their reliability practices, and mentor associate engineers.
- Share field learnings with product and engineering teams.
Required Qualifications
- Substantial SRE experience with real ownership of production reliability on AWS.
- Experience running Kubernetes and AWS services at scale.
- A track record of defining and managing SLIs, SLOs, and error budgets.
- Strong observability practice.
- Proven incident leadership and post-mortem facilitation.
- Strong automation skills (Python, Go, or Bash).
- Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
- AWS certification at Associate level as a minimum.
Preferred Qualifications
- Owning CI/CD pipelines and release automation.
- Writing and maintaining Terraform modules.
- GitOps workflows and Helm.
- Designing telemetry collection across services.
- Incident-management tooling such as PagerDuty.
- Hands-on experience integrating AIOps or AI SRE tooling in