01 Sep
|
NCR Voyix
|
Chennai
Reliability & Service Availability
- Monitor enterprise applications, infrastructure, cloud platforms, and customer-facing services.
- Ensure platform availability, performance, and service health within established SLAs and SLOs.
- Identify service degradation trends and proactively address reliability risks.
- Drive operational improvements to reduce incidents and improve system resiliency.
- Establish and track reliability metrics, KPIs, and service health indicators.
- Develop preventative measures to reduce future service disruptions.
Observability & Monitoring
- Configure and maintain monitoring platforms including:
- AppDynamics
- Dynatrace
- Datadog
- Splunk
- Azure Monitor
- Google Cloud Operations
- ServiceNow Event Management
- Develop dashboards, alerts, and health monitoring solutions.
- Reduce alert fatigue through alert tuning and optimization.
Automation & Engineering
- Automate operational processes and repetitive tasks.
- Develop scripts and tooling using:
- PowerShell
- Python
- Bash
- APIs
- Improve operational efficiency through self-healing and automated remediation capabilities.
- Support infrastructure-as-code and reliability engineering initiatives.
Operational Readiness
- Participate in Operational Readiness Reviews (ORR).
- Validate monitoring, alerting, runbooks, and support procedures prior to production go-live.
- Ensure escalation paths and support models are documented and operational.
- Review application deployments for supportability and operational risks.
Cloud & Infrastructure Support
- Support hybrid environments across:
- Azure
- Google Cloud Platform (GCP)
- AWS
- VMware
- Analyze application, network, database, and infrastructure performance issues.
- Work with engineering teams to optimize platform stability and scalability.
Governance & Reporting
- Produce incident reports, service health updates, operational reviews, and executive summaries.
- Maintain operational documentation, runbooks, and knowledge articles.
- Track service performance metrics and reliability improvements.
- Support audit and compliance initiatives as required.
Required Qualifications
- 6+ years of experience in:
- Site Reliability Engineering
- Production Support
- Systems Engineering
- DevOps
- Network Operations Center (NOC)
- Command Center Operations
- Experience supporting mission-critical production environments.
- Strong understanding of
- Incident Management
- Problem Management
- Change Management
- Service Level Management
- Experience with ServiceNow or similar ITSM platforms.
Technical Skills
Operating Systems
- Windows Server
- Linux/Unix
Cloud Platforms
- Microsoft Azure
- Google Cloud Platform (GCP)
- Amazon Web Services (AWS)
Monitoring & Observability
- AppDynamics
- Datadog
- Dynatrace
- Recent Relic
- Splunk
- Azure Monitor
- ServiceNow Event Management
Preferred Qualifications
- Experience working in a Global Command Center environment.
- AWS, Azure, or GCP certifications.
- Experience supporting retail, hospitality, payments, or enterprise SaaS platforms.
- Experience with CI/CD pipelines and DevOps practices.
- Knowledge of SRE concepts including:
- SLI/SLO/SLA management
- Error budgets
- Chaos testing
- Resiliency engineering
Key Competencies
- Critical Incident Leadership
- Technical Troubleshooting
- Problem Solving
- Customer Focus
- Operational Excellence
- Communication Skills
- Executive Presence
- Collaboration
- Continuous Improvement
- Decision Making Under Pressure
Success Metrics The GCC Site Reliability Engineer will be measured on:
- Service availability and uptime
- Incident response times
- Mean Time to Detect (MTTD)
- Mean Time to Restore (MTTR)
- Reduction in recurring incidents
- Monitoring effectiveness
- Automation adoption
- Operational readiness compliance
- Customer impact reduction
- Service reliability improvements
Work Environment
- 24x7 operational support organization.
- Participation in on-call and major incident rotations.
- Collaboration with global teams across multiple regions.
- Hybrid cloud and enterprise production environments.
- Fast-paced, mission-critical operational setting.
📌 Site Recovery Manager (Chennai)
🏢 NCR Voyix
📍 Chennai