Job Summary
A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.
Key Responsibilities
Own the reliability of production systems for one or more enterprise customers on AWS.
Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
Lead the response to responsible (P0/P1) incidents and run blameless post-mortems that lead to fixes.
Lead within the on-call rotation.
Operate and optimise Kubernetes clusters and AWS services as workloads grow.
Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
Advise customers on improving their reliability practices, and mentor associate engineers.
Share field learnings with product and engineering teams.
Required Qualifications
Substantial SRE experience with real ownership of production reliability on AWS.
Experience running Kubernetes and AWS services at scale.
A track record of defining and managing SLIs, SLOs, and error budgets.
Solid observability practice.
Proven incident leadership and post-mortem facilitation.
Solid automation skills (Python, Go, or Bash).
Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
AWS certification at Associate level as a minimum.
Preferred Qualifications
Owning CI/CD pipelines and release automation.
Writing and maintaining Terraform modules.
GitOps workflows and Helm.
Designing telemetry collection across services.
Incident-management tooling such as PagerDuty.
Hands-on experience integrating AIOps or AI SRE tooling into production operations.
Familiarity with observabi