09 Sep
|
Infinite
|
Hyderabad
09 Sep
Infinite
Hyderabad
Key Responsibilities
Crisis Bridge Command & Service Restoration (P1 / P2 / P3 Impacting Incidents)
- Chair and command live incident bridges for service-impacting outages (P1 / P2 / P3), driving rapid restoration with cross-functional technical teams (Network, Cloud, Security, Infrastructure, DevOps).
- Enforce structured triage methodologies to quickly isolate root triggers and drive mean time to restore (MTTR) metrics downward during high-priority disruptions.
- Evaluate risk and approve Emergency Change Requests (eCAB) required to restore compromised operational services.
Proactive Incident Management & Prevention (P4 Incidents)
- Lead proactive management workflows (P4) by monitoring health alerts, performance degradations, and system anomalies to resolve vulnerabilities before they impact business operations.
- Partner with monitoring and engineering teams to establish automated triggers and actionable playbooks for early-stage P4 events.
Executive & Stakeholder Communications
- Draft and distribute explicit, real-time broadcast communications, status dashboards, and executive briefings under high pressure for active P1 / P2 / P3 events.
- Translate highly complex technical diagnostics into plain-language business impact statements for C-suite leaders, business unit directors, and key stakeholders.
Process Management & Governance
- Ensure strict operational adherence to ITIL v3/ITIL 4 Incident, Problem, and Change Management frameworks across global operations.
- Partner with Problem Management to transition resolved major incidents into formal Post-Incident Reviews (PIR/RCA) using 5-Why and Fishbone analysis.
Operational Excellence & Continuous Improvement
- Monitor and analyze operational metrics (MTTD, MTTR, escalation efficiency, P4-to-P1 prevention rates) to identify operational friction points.
- Refine emergency response playbooks, communication matrices, and automated alert notification workflows.
Qualifications & Skillset Requirements
Primary / Required Qualifications
- Experience: 610 years in IT Service Management, with at least 4+ years dedicated to leading Major Incident Management (MIM) command bridges for service-impacting outages (P1 / P2 / P3) in large-scale enterprise environments.
- Incident Command & Proactive Triage: Proven ability to direct multi-disciplinary teams under high stress during live outages, as well as a track record of driving proactive (P4) fault isolation to prevent escalation.
- Communication: Exceptional verbal and written command of English; proven track record of authoring executive-level outage summaries and client notifications under tight SLAs.
- Technical Domain Awareness:Solid operational understanding of enterprise IP networks, routing/switching, 5G/LTE transport, and multi-cloud environments (AWS/Azure).
- Framework Mastery: Deep expertise in ITIL framework principles (ITIL v3 or ITIL 4 Foundation certification preferred).
Secondary / Preferred Qualifications
- Tooling Proficiency: Hands-on experience with ITSM platforms (ServiceNow, Jira Service Management) and alerting tools (xMatters, PagerDuty, Slack).
- Observability: Ability to navigate telemetry and monitoring dashboards (Grafana, AWS Health Dashboard, Splunk) to identify proactive P4.
📌 Incident Management Specialist (Hyderabad)
🏢 Infinite
📍 Hyderabad