Observability & Monitoring Lead - DevOps (India)

Observability & Monitoring Lead - DevOps (India)

06 Aug
|
Tidyhire
|
India

06 Aug

Tidyhire

India

Key Responsibilities:

- Lead the design, implementation, and governance of enterprise observability and monitoring solutions.
- Build and maintain monitoring platforms using Prometheus, Grafana, Azure Monitor, Application Insights, Splunk, and OpenTelemetry.
- Develop dashboards, alerts, and SLO/SLI metrics to proactively monitor application and infrastructure health.
- Monitor and optimize the performance, availability, and reliability of cloud platforms, Kubernetes clusters, and business-critical applications.
- Lead major incident response, root cause analysis (RCA), and continuous service improvement initiatives.
- Implement centralized logging, distributed tracing, and end-to-end observability across cloud-native environments.
- Automate monitoring deployment, alerting, and operational workflows using Terraform, Ansible, Python, PowerShell, or Bash.
- Collaborate with Platform Engineering, DevOps, SRE, Security, and Application teams to establish monitoring standards and best practices.
- Drive capacity planning, performance tuning, and cost optimization for monitoring platforms.
- Mentor engineers on observability practices and foster a culture of reliability, automation, and operational excellence.

Key Skills and Experience:

- 1216 years of overall IT experience, with at least 6+ years of hands-on experience in Observability, Monitoring, Site Reliability Engineering (SRE), or Platform Engineering.
- Proven experience in designing, implementing,



and managing enterprise observability and monitoring solutions for large-scale production environments.
- Robust hands-on expertise with Prometheus, Grafana, Azure Monitor, Application Insights, OpenTelemetry, and centralized logging platforms such as Splunk or ELK.
- Extensive experience monitoring Azure cloud infrastructure, Kubernetes (AKS), containerized applications, distributed systems, and cloud-native platforms.
- Hands-on experience in building dashboards, alerts, metrics, logs, and distributed tracing to ensure end-to-end visibility and proactive monitoring.
- Strong understanding of Site Reliability Engineering (SRE) principles, including SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis (RCA), performance optimization, and service reliability.
- Experience supporting enterprise production environments, managing major incidents, and driving high availability, operational excellence, and continuous service improvement.
- Proficiency in automation and Infrastructure as Code (IaC) using Terraform, Ansible, Python, PowerShell, or Bash.
- Strong troubleshooting, analytical, communication, stakeholder management, and cross-functional collaboration skills.

Preferred Certifications:

- Microsoft Certified: Azure Administrator (AZ-104)
- Microsoft Certified: Azure Solutions Architect (AZ-305)
- Certified Kubernetes Administrator (CKA)
- Grafana Certified Professional
- Prometheus Certified Associate (PCA)
- ITIL Foundation
- SRE Foundation / SRE Practitioner

📌 Observability & Monitoring Lead - DevOps (India)
🏢 Tidyhire
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: observability & monitoring lead - devops (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: observability & monitoring lead - devops (india) / india