10 Aug
|
Capgemini
|
Delhi
We are hiring a hands-on Service Delivery Manager to run day-to-day production support for our Big Data and data integration platforms. You will lead a support pod of L1/L2/L3 engineers, act as incident commander for critical outages, own SLA delivery, and drive the automation agenda that removes recurring manual work.
Key Responsibilities
Incident & Problem Management
- Act as incident commander for P1/P2 production outages drive triage bridges, coordinate cross-team resolution, and manage business communication until service restoration.
- Own and deliver Root Cause Analysis (RCA) documentation within agreed SLAs; track corrective and preventive actions to closure.
- Run problem management for recurring issues; convert known errors into permanent fixes rather than repeated workarounds.
SLA & Metric Governance
- Track, report, and enforce response and resolution SLAs across the ticket queue; manage backlog, ageing, and breach risk proactively.
- Publish weekly and monthly KPI packs MTTR, MTTD, SLA compliance, ticket volume and trend, availability, automation savings to leadership.
- Own the operational review cadence with business and technology stakeholders.
People & Team Leadership
- Manage 16/7 or 24/7 rotational support rosters, shift handovers, on-call schedules, and holiday coverage.
- Mentor and coach support engineers; build technical depth and runbook discipline across the team.
- Conduct performance reviews, goal setting, and development planning; participate actively in interviewing, hiring, and onboarding.
Automation & Continuous Improvement
- Identify high-frequency,
high-effort manual tasks and replace them with automation scripts, self-healing routines, and auto-remediation.
- Improve monitoring and alerting frameworks to shift from reactive to predictive detection; reduce alert noise and false positives.
- Maintain and continually improve runbooks, knowledge base articles, and standard operating procedures.
Cross-Team Collaboration
- Partner with software engineering, DevOps, and product teams to review upcoming deployments and assess release readiness.
- Represent support in CAB and change approval forums; enforce operational readiness criteria before production releases.
- Feed stability pain points, defect trends, and supportability gaps back into the development backlog.
Must Have
- Hands-on experience supporting Hadoop ecosystem components (HDFS, Hive, Spark, YARN, Sqoop, Oozie).
- Working experience with IBM InfoSphere DataStage — job design, scheduling, failure triage, and performance troubleshooting.
- Solid Splunk skills — building queries, dashboards, and alerts for log analysis and monitoring.
- Strong SQL for data validation, reconciliation, and issue investigation.
- Scripting/programming in Python, Java, or Shell for automation and tooling.
- ITIL-based Incident, Problem, and Change Management; ServiceNow or JIRA.
- Linux/Unix fundamentals and job scheduling tools (Control-M, Autosys, Airflow).
Good to Have
- Cloud data platforms (AWS, Azure, GCP), Databricks, or Snowflake.
- CI/CD and DevOps tooling exposure (Jenkins, Git, Ansible).
- ITIL v4 Foundation certification.
📌 Service Delivery Manager (Delhi)
🏢 Capgemini
📍 Delhi