Senior Observability Engineer - Grafana & Prometheus (Bengaluru)

Senior Observability Engineer - Grafana & Prometheus (Bengaluru)

14 Aug
|
Zensar Technologies
|
Bengaluru

14 Aug

Zensar Technologies

Bengaluru

Job Summary

Zensar Technologies is looking for a highly skilled Senior Observability Engineer to design, implement, and manage enterprise-scale monitoring and observability solutions for a large AWS cloud modernization program. The ideal candidate will have strong hands-on expertise in Grafana, Prometheus, Monitoring & Alerting, Dashboard Development, Metrics Collection, and Troubleshooting, along with exposure to AWS observability services.

The role involves building a centralized observability platform to provide end-to-end visibility, reliability, performance monitoring, proactive alerting, and faster incident resolution across large-scale enterprise applications and integration environments. Key Responsibilities

Observability Platform Engineering

Design and implement enterprise observability solutions using Grafana and Prometheus.

Configure and manage monitoring platforms for metrics collection, visualization, and reporting.

Build AWS-native observability capabilities using CloudWatch, X-Ray, AWS Managed Grafana, and OpenTelemetry.

Implement monitoring and instrumentation across distributed applications and services.

Ensure comprehensive visibility into application, infrastructure, and platform performance.

Monitoring & Dashboard Development

Develop and maintain operational dashboards for infrastructure, applications, and business KPIs.

Create real-time dashboards for performance monitoring, SLA tracking, and operational health checks.

Design monitoring frameworks to track system utilization, availability, and service reliability.

Provide actionable insights through visualization and performance analytics.

Alerting & Reliability Engineering

Configure alerting rules, thresholds, and notification mechanisms.

Implement proactive monitoring and intelligent alerting capabilities.

Support service reliability objectives by identifying and mitigating performance issues before business impact.





Drive observability best practices for incident prevention and operational excellence.

Troubleshooting & Incident Support

Perform root cause analysis using metrics, logs, traces, and dashboards.

Support production troubleshooting and incident investigations.

Work closely with engineering, operations, and platform teams to resolve critical issues.

Improve mean time to detect (MTTD) and mean time to resolve (MTTR).

Governance & Best Practices

Define observability standards, monitoring frameworks, and dashboard templates.

Establish logging, tracing, and correlation standards across systems.

Develop monitoring runbooks and operational documentation.

Drive continuous improvement initiatives across observability and reliability practices.

Required

Skills

Strong hands-on experience with Grafana and Prometheus.

Experience in Monitoring, Observability, Alerting, Dashboard Development, and Metrics Collection.

Expertise in troubleshooting enterprise applications and production environments.

Strong understanding of

Monitoring & Alerting

Logging & Metrics Management

Distributed Tracing

Site Reliability Engineering (SRE) Concepts

Experience with AWS Observability Services:

CloudWatch

AWS Managed Grafana

AWS X-Ray

Excellent analytical and problem-solving skills.

Experience supporting large-scale enterprise applications and cloud environments. Positive to Have

OpenTelemetry (OTEL) instrumentation.

Kubernetes / EKS monitoring.

SRE practices and reliability engineering.

Experience with enterprise integration platforms and cloud modernization initiatives.

AI-driven observability, anomaly detection, and predictive monitoring.

Exposure to high-volume transaction processing environments.

Preferred

Qualifications

Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipline.

8-14 years of overall IT experience.

Proven experience in Observability Engineering, Site Reliability Engineering (SRE), Monitoring Engineering, or Platform Operations.

Experience working in AWS cloud environments.

📌 Senior Observability Engineer - Grafana & Prometheus (Bengaluru)
🏢 Zensar Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior observability engineer - grafana & prometheus (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior observability engineer - grafana & prometheus (bengaluru) / bengaluru