09 Sep
|
Infinite
|
Hyderabad
09 Sep
Infinite
Hyderabad
Key Responsibilities
Crisis Bridge Command & Service Restoration (P1 / P2 / P3 Impacting Incidents)
* Chair and command live incident bridges for service-impacting outages (P1 / P2 / P3), driving rapid restoration with cross-functional technical teams (Network, Cloud, Security, Infrastructure, DevOps).
* Enforce structured triage methodologies to quickly isolate root triggers and drive mean time to restore (MTTR) metrics downward during high-priority disruptions.
* Evaluate risk and approve Emergency Change Requests (eCAB) required to restore compromised operational services.
Proactive Incident Management & Prevention (P4 Incidents)
* Lead proactive management workflows (P4) by monitoring health alerts, performance degradations, and system anomalies to resolve vulnerabilities before they impact business operations.
* Partner with monitoring and engineering teams to establish automated triggers and actionable playbooks for early-stage P4 events.
Executive & Stakeholder Communications
* Draft and distribute explicit, real-time broadcast communications, status dashboards, and executive briefings under high pressure for active P1 / P2 / P3 events.
* Translate highly complex technical diagnostics into plain-language business impact statements for C-suite leaders, business unit directors, and key stakeholders.
Process Management & Governance
* Ensure strict operational adherence to ITIL v3/ITIL 4 Incident, Problem, and Change Management frameworks across global operations.
* Partner with Problem Management to transition resolved major incidents into formal Post-Incident Reviews (PIR/RCA) using 5-Why and Fishbone analysis.
Operational Excellence & Continuous Improvement
* Monitor and analyze operational metrics (MTTD, MTTR, escalation efficiency, P4-to-P1 prevention rates) to identify operational friction points.
* Refine emergency response playbooks, communication matrices, and automated alert notification workflows.
Qualifications & Skillset Requirements
Primary / Required Qualifications
* Experience: 610 years in IT Service Management, with at least 4+ years dedicated to leading Major Incident Management (MIM) command bridges for service-impacting outages (P1 / P2 / P3) in large-scale enterprise environments.
* Incident Command & Proactive Triage: Proven ability to direct multi-disciplinary teams under high stress during live outages, as well as a track record of driving proactive (P4) fault isolation to prevent escalation.
* Communication: Exceptional verbal and written command of English; proven track record of authoring executive-level outage summaries and client notifications under tight SLAs.
* Technical Domain Awareness:Solid operational understanding of enterprise IP networks, routing/switching, 5G/LTE transport, and multi-cloud environments (AWS/Azure).
* Framework Mastery: Deep expertise in ITIL framework principles (ITIL v3 or ITIL 4 Foundation certification preferred).
Secondary / Preferred Qualifications
* Tooling Proficiency: Hands-on experience with ITSM platforms (ServiceNow, Jira Service Management) and alerting tools (xMatters, PagerDuty, Slack).
* Observability: Ability to navigate telemetry and monitoring dashboards (Grafana, AWS Health Dashboard, Splunk) to identify proactive P4.
📌 Incident Management Specialist (Hyderabad)
🏢 Infinite
📍 Hyderabad