We are looking for an experienced Site Reliability Engineer with experience in Lead, cloud infrastructure, platform reliability, automation, observability, and production support.
Key Responsibilities
- Manage and improve reliability of Azure cloud and Kubernetes platforms.
- Define and monitor SLIs, SLOs, and error budgets.
- Build and maintain CI/CD pipelines using Jenkins and Azure DevOps.
- Manage containerized applications using Docker, Kubernetes, and Helm.
- Automate infrastructure using Terraform and Ansible.
- Implement monitoring and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
- Handle L2/L3 production support, incident management, troubleshooting, and RCA.
- Support PostgreSQL, Redis, and RabbitMQ environments.
- Drive automation, performance, scalability, and reliability improvements.