Seeking a Senior Infrastructure Engineer with strong Azure, Kubernetes, automation and observability experience. The role focuses on infrastructure operations, self-healing automation, monitoring and production support for an AIOps programme.
- Manage production, canary and non-production environments
- Implement self-healing for certificates, queues, ADF and Kubernetes
- Integrate logs, metrics, alerts and events for RCA
- Implement rollback, audit and safe automation controls
- Manage certificates, queues and Kubernetes platforms
- Handle production monitoring, incident triage and hyper-care
- Maintain runbooks, TSGs and documentation
- Collaborate with infrastructure, security and engineering teams