02 Sep
|
NCR Voyix
|
Chennai
Position Overview
We are seeking an experienced Observability Platform Manager to lead the strategy, implementation, and continuous evolution of enterprise observability capabilities across cloud and hybrid environments. The ideal candidate has recent hands-on experience implementing observability tools at scale, a strong foundation in cloud engineering, and the leadership skills to drive platform adoption, standardization, and operational excellence across engineering teams. Familiarity with AI applications and AI-enabled observability capabilities is a strong plus.
Key Responsibilities Platform Strategy Leadership
- Define and execute the vision, roadmap, and operating model for the enterprise observability platform, including logs, metrics, traces, dashboards, and alerting.
- Lead and mentor engineers responsible for building, operating, and improving observability capabilities across business-critical platforms and applications.
- Establish platform standards, governance, onboarding patterns, and success measures to drive consistent adoption at scale.
- Partner with SRE, platform engineering, cloud engineering, infrastructure, and application teams to align observability strategy with reliability and business goals.
Observability Platform Implementation
- Lead recent and large-scale implementations of observability tools and platforms across multi-team or enterprise environments.
- Evaluate, select, and optimize observability tooling such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, Elasticsearch, or equivalent solutions based on scale, cost, and business needs.
- Drive implementation of telemetry standards, ingestion pipelines, access models, dashboard frameworks, and alerting practices that improve signal quality and reduce operational noise.
- Manage vendor relationships, platform lifecycle decisions, and cost/performance tradeoffs for observability capabilities.
Cloud Engineering Reliability
- Bring strong cloud engineering experience across AWS, Azure, or GCP,
with an understanding of cloud-native architectures, resilience patterns, and scalable operations.
- Guide teams on instrumentation and monitoring for microservices, containers, Kubernetes, serverless workloads, and distributed systems using modern telemetry standards such as OpenTelemetry.
- Partner with engineering teams to embed observability into platform design, CI/CD pipelines, incident response, and service reliability practices.
Operational Excellence Enablement
- Establish service-level indicators, objectives, reporting, and operational health reviews to improve platform reliability and engineering outcomes.
- Develop enablement programs, best practices, and self-service patterns that help teams adopt observability consistently and effectively.
- Drive continuous improvement in incident detection, triage, root-cause analysis, and post-incident learning through better telemetry and platform workflows.
Automation AI Applications
- Promote automation-first practices for instrumentation, alert tuning, dashboard provisioning, and operational workflows.
- Identify opportunities to apply AI-enabled capabilities such as anomaly detection, event correlation, intelligent alerting, and operational insights.
- Familiarity with AI applications, AI platforms, or AI-supported engineering workflows is a plus.
Required Qualifications
- 10+ years of experience in observability, SRE, platform engineering, cloud engineering, or related disciplines, including recent experience implementing observability tools at scale.
- Proven experience leading or managing observability platforms, programs,
or engineering teams in complex enterprise environments.
- Hands-on experience with enterprise observability and monitoring tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, Elasticsearch, or comparable platforms.
- Strong cloud engineering experience with AWS, Azure, or GCP, including cloud-native services, architecture patterns, and operational best practices.
- Knowledge of distributed systems, microservices, containers, Kubernetes, CI/CD, and infrastructure-as-code practices.
- Familiarity with OpenTelemetry , distributed tracing, service health models, and telemetry data management.
- Ability to define meaningful KPIs, SLIs/SLOs, alerting strategies, and executive-ready reporting that tie technical health to business outcomes.
- Strong collaboration and communication skills, with the ability to influence stakeholders across engineering, operations, and leadership teams.
Preferred Qualifications
- Familiarity with AI applications, AI engineering workflows, or AI-enhanced observability capabilities.
- Knowledge of container ecosystems and orchestration platforms (Kubernetes, AKS/EKS/GKE).
- Experience working with event-driven architectures and microservices environments.
- Robust scripting or programming skills (Python, PowerShell, Bash, etc.).
- Relevant certifications (e.g., Splunk Architect, Dynatrace Professional, Cloud certifications).
Soft Skills
- Excellent communication and stakeholder management skills.
- Ability to lead technical strategy and influence architectural decisions.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Adaptability and curiosity about new technologies and evolving observability trends.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Observability Platform Manager (Chennai)
🏢 NCR Voyix
📍 Chennai