Site Reliability Engineer/ Expert/ Specialist (New Delhi)

Site Reliability Engineer/ Expert/ Specialist (New Delhi)

29 Sep
|
SITA
|
New Delhi

29 Sep

SITA

New Delhi

Job Summary

Welcome to SITA. At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. You'll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We don't just move the world forward-we're proud to be recognized as a Outstanding Place to Work by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. Are you ready to love your job? The adventure begins right here, with you, at SITA.

Purpose. Ensure high product performance, reliability, and stability by proactively supporting products, resolving root causes of incidents, and implementing improvements that prevent recurrence. This role focuses on event management, automation, service deployment, and operational integration to improve efficiency and collaboration across service operations.

Responsibilities

- Build and maintain reliable support systems to ensure high availability and strong product performance.
- Manage complex operational cases, incidents, and root cause analysis to deliver permanent fixes.
- Define and maintain event catalogs, alerts, thresholds, and remediation actions.
- Implement automation for provisioning, monitoring, deployment, self-healing, and recovery.
- Collaborate with Product, Engineering, Service Architecture, and Operations teams to improve service readiness, availability, and performance.
- Support customer success initiatives through reporting, documentation, communication materials, and process improvements.
- Contribute to knowledge management resources such as FAQs, training materials, and operational guidance.




- Apply data governance standards, monitor data quality, and act as a subject matter expert for data-related queries.

Qualifications

- Bachelor's degree in computer science, Information Technology, Engineering, or a related field.
- 5+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager.
- Proven experience managing high-availability systems and ensuring operational reliability with Azure/Windows Environments.
- Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions.
- Hands-on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC).
- Strong background in collaborating with cross-functional teams (Development, Operations, Engineering, etc.) to improve operational processes and service delivery.
- Experience managing deployments, conducting risk assessments, and optimizing event and problem management processes.
- Familiarity with cloud technologies, containerization, and scalable architectures, including zero-downtime deployment strategies.
- Windows Server OS expertise (Active Directory, Group Policy, DNS, DHCP).
- Solid problem management and troubleshooting skills.
- Working knowledge of Unix/Linux (Red Hat).
- PowerShell scripting.
- Azure and AWS experience.
- AKS and on-prem Kubernetes knowledge.
- Automation experience,



including CI/CD pipelines and exposure to Terraform.
- VMware expertise (VCF, VCD).
- Experience with SAN and NAS storage management.
- Backup solutions experience with Veeam and Veritas.
- Networking knowledge at CCNA level.
- Databases: SQL and MongoDB for restore operations and performance tuning.
- Observability and monitoring: Prometheus, Grafana, Dynatrace, Azure Monitor, and AWS CloudWatch.
- Experience with enterprise monitoring and observability platforms; designing proactive monitoring, alerting, and capacity management to improve reliability and reduce incidents; knowledge of log management, metrics, dashboards, distributed tracing, and root cause analysis; defining and monitoring SLIs, SLOs, and SLAs to support SRE practices; focus on operational excellence and platform stability.

Core Competencies

- Problem Management: Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions.
- Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions.
- Monitor the effectiveness of problem resolution activities and provide regular reporting to ensure continuous improvement.

Certifications and Additional Qualifications

- Relevant certifications: VMware, AWS/Azure; MCSE; RHSA certifications are highly recommended.
- CompTIA Security+ or Certified Kubernetes Administrator (AKS or CKA).
- Certifications in cloud platforms (AWS, Azure, Google Cloud) or DevOps methodologies (e.g. Certified DevOps Professional).

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Site Reliability Engineer/ Expert/ Specialist (New Delhi)
🏢 SITA
📍 New Delhi

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/ expert/ specialist (new delhi) / new delhi

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/ expert/ specialist (new delhi) / new delhi