30 Jul
|
NCR Voyix
|
Telangana
30 Jul
NCR Voyix
Telangana
Key Responsibilities
Design and implement enterprise observability solutions across Azure, GCP, Kubernetes, and hybrid environments.
Develop monitoring, logging, tracing, and telemetry standards using industry best practices.
Build enterprise dashboards that provide real-time infrastructure, application, and customer health visibility.
Define and implement SLIs, SLOs, error budgets, and operational health metrics.
Improve proactive detection, alert quality, and incident response through automation and intelligent alerting.
Integrate observability with ServiceNow, automation platforms, and operational workflows.
Partner with Product Engineering, Infrastructure, Security, and Operations teams to improve platform reliability and operational readiness.
Support enterprise initiatives involving AI-driven observability, event correlation, and operational analytics.
Required Skills
Site Reliability Engineering (SRE)
Kubernetes (AKS/GKE)
Azure and Google Cloud Platform
Grafana, Datadog, Prometheus, OpenTelemetry (or similar)
Monitoring, logging, distributed tracing, and telemetry
Infrastructure as Code (Terraform)
Python, Go, or PowerShell automation
CI/CD and DevOps practices
📌 Senior Site Reliability Engineer (Telangana)
🏢 NCR Voyix
📍 Telangana