Monitoring & Observability Delivery Manager (India)

Monitoring & Observability Delivery Manager (India)

07 Aug
|
Relevance Labs
|
India

07 Aug

Relevance Labs

India

Job Title: Monitoring & Observability Delivery Manager

Experience

12–18 Years

Location

Bangalore / Hybrid

Role Summary

We are looking for an experienced Monitoring & Observability Delivery Manager to lead the delivery, operations, and continuous improvement of enterprise observability platforms across cloud and on-premises environments. The ideal candidate should possess strong expertise in logs, metrics, traces, dashboards, alerting, incident management, and SRE practices, along with proven experience in managing large delivery teams and customer engagements.

This role requires a combination of technical expertise, delivery leadership, stakeholder management, and operational excellence to ensure high platform availability, proactive monitoring, and rapid incident resolution.

Key Responsibilities

Delivery Leadership

- Lead end-to-end delivery of Monitoring & Observability services across multiple customer environments.

- Manage large teams of engineers, technical leads, and architects.

- Drive project execution, planning, governance, and customer communication.

- Ensure SLA, KPI, and operational excellence targets are consistently achieved.

- Manage production support, incident management, problem management, and change management processes.

- Present service health, operational metrics, and executive dashboards to customer leadership.

Monitoring & Observability

- Design and implement enterprise monitoring strategies.

- Build observability solutions using logs, metrics, traces, and distributed tracing.

- Define monitoring standards, alerting strategies, dashboards, and operational KPIs.

- Drive observability maturity across applications, infrastructure, cloud, and Kubernetes platforms.

- Reduce alert noise through alert tuning and automation.

Log Analytics

Strong understanding of:

- Application Logs

- Infrastructure Logs

- Kubernetes Logs

- Linux & Windows Logs





- Database Logs

- Middleware Logs

- API Gateway Logs

- Audit Logs

- Security Logs

Ability to

- Analyze log patterns

- Troubleshoot production incidents

- Identify recurring failures

- Perform Root Cause Analysis (RCA)

- Build centralized log analytics solutions

Metrics & Performance Monitoring

Experience in monitoring

- CPU

- Memory

- Disk

- Network

- JVM

- Containers

- Kubernetes

- Database Performance

- API Performance

- Application Availability

- Response Time

- Throughput

- Error Rates

- Capacity Planning

- Infrastructure Health

Should be capable of defining

- Golden Signals

- RED Metrics

- USE Metrics

- SLI/SLO/SLA

- Business KPIs

Cloud & Platform Monitoring

Hands-on knowledge of

- AWS CloudWatch

- Azure Monitor

- Google Cloud Operations Suite

- Kubernetes (EKS/AKS/OpenShift)

- Docker

- Linux

- Windows Servers

- VMware

- Databases

- Middleware

Observability Tools

Hands-on experience with one or more of the following

- Dynatrace

- Datadog

- Splunk

- Elastic (ELK)

- Grafana

- Prometheus

- Current Relic

- AppDynamics

- OpenTelemetry

- Jaeger

- Zipkin

Incident & Problem Management

- Lead P1/P2 production incidents.

- Drive Root Cause Analysis (RCA).

- Identify recurring issues and preventive actions.

- Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).

- Implement proactive monitoring and automation.

Automation

Experience with

- Python

- Shell Scripting

- PowerShell

- Terraform

- Ansible

- Jenkins





- GitHub Actions

Automation of

- Monitoring deployment

- Alert creation

- Dashboard provisioning

- Health checks

- Incident remediation

Customer & Stakeholder Management

- Work closely with customer leadership, architects, infrastructure teams, and application teams.

- Present weekly/monthly operational reviews.

- Manage escalations and executive communications.

- Define roadmap for observability transformation.

Required Skills

- 12+ years of IT experience.

- 5+ years managing Monitoring & Observability platforms.

- Strong understanding of logs, metrics, traces, and distributed systems.

- Experience managing cloud-native environments and Kubernetes.

- Strong troubleshooting and production support expertise.

- Excellent communication, stakeholder management, and leadership skills.

- Experience managing teams of 20+ engineers.

Preferred Qualifications

- AWS/Azure/GCP Certification

- Dynatrace Professional Certification

- Datadog Certification

- Splunk Certification

- Kubernetes (CKA/CKAD)

- ITIL Foundation

- SRE or DevOps certifications

Success Metrics

- SLA Compliance (>99%)

- Platform Availability

- Reduced MTTD and MTTR

- Alert Noise Reduction

- Dashboard Adoption

- Incident Reduction

- Automation Coverage

- Customer Satisfaction (CSAT)

- Team Utilization and Delivery Excellence

Ideal Candidate Profile

The ideal candidate is a technically strong delivery leader who can interpret logs, metrics, and traces to quickly identify production issues, lead cross-functional incident response, and build scalable observability solutions. They should have experience delivering enterprise monitoring services, mentoring technical teams, engaging with customers, and driving continuous improvements through automation and observability best practices.

📌 Monitoring & Observability Delivery Manager (India)
🏢 Relevance Labs
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: monitoring & observability delivery manager (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: monitoring & observability delivery manager (india) / india