- Ensure high availability, scalability, and reliability of production systems
- Design and implement monitoring, alerting, and observability solutions
- Lead incident response and post-mortem analysis for outages
- Automate operational processes to reduce manual toil
- Define and track SLIs/SLOs/error budgets for critical services
- Collaborate with development teams on reliability best practices
Required Skills:
- Strong experience with cloud infrastructure (AWS/GCP/Azure)
- Hands-on experience with Kubernetes and containerization
- Experience with monitoring/observability tools (Prometheus, Grafana, Datadog)
- Strong scripting/programming skills (Python/Go/Bash)
- Understanding of incident management and on-call practices
- Robust troubleshooting and systems thinking skills
Good to Have:
- Experience with chaos engineering practices
- Knowledge of Terraform/Infrastructure-as-Code
- Experience running large-scale distributed systems