Leads reliability improvements across applications, platforms, and cloud systems. Drives automation, enhances observability, optimizes performance, and conducts root-cause analysis. Partners with engineering teams to reduce toil, improve operational maturity, and strengthen service resilience.
Cloud Platform : Azure
Minimum Degree Required: Bachelors
Preferred candidate profile
Required / Mandatory Knowledge/Skills: (character count limit 5000) PLEASE ONLY USE THIS FIELD IF THIS IS A MUST HAVE SKILL FOR APPLICANT
Solid understanding of SRE practices including SLIs/SLOs, error budgets, service health, and operational KPIs
Ability to automate operational tasks using Python, Shell, PowerShell, Go, or similar languages
Experience improving alerting systems, reducing noise,
and refining observability instrumentation
Proficiency with cloud platforms and core services (compute, storage, networking, serverless)
Experience executing root-cause analysis and problem management
Ability to lead incident response and coordinate cross-team troubleshooting
Experience identifying systemic reliability gaps and proposing engineering solutions
Ability to design performance tests, validate reliability risks, and assess scalability
Robust communication skills for partnering with development, operations, and leadership