Required Skills:
- 7–12 years of experience in Cloud Ops / SRE / NOC environments (24x7 operations)
- Robust expertise in Azure Infrastructure (VMs, Networking, Storage)
- Hands-on experience with Azure Kubernetes Service (AKS), Kubernetes, Docker
- Strong experience with monitoring and observability tools (Datadog, Azure Monitor)
- Proven experience in Incident Management / Major Incident Handling, Monthly reporting
- Experience with Infrastructure as Code (Terraform, ARM templates, Helm)
- Scripting skills in Power Shell, Python, or Bash
- Experience with Service Now (Incident, Problem, Change modules and dashboards)
- Good understanding of distributed systems and cloud-native architecture
- Excellent communication, leadership, and problem-solving skills
Originally posted on Himalayas
📌 Lead Software Engineer, Cloud Site Reliability (SRE) (India)
🏢 Icertis
📍 India