1. Experience in supporting large-scale distributed systems
2. Expertise in Application and batch monitoring and troubleshooting
3. Expertise in Incident and problem management
4. Expertise in Production deployments (Blue-Green and Canary deployments) using CI/CD pipelines
5. Troubleshooting deployment failures (L1)
6. Experience supporting production maintenance activities such as DR, certificate renewals, and password renewals
7. Solid understanding of SLI, SLO, SLA, error budgets, and burn rate alerts
8. Expertise in APM tools, including writing log queries and building dashboards
9. Basic Shell and Python scripting skills
10. Linux/Windows server administration and troubleshooting (L1)
11. Cloud operations experience (L1)
12. Network, load balancer, and database failure troubleshooting (L1)