Role & responsibilities
Incident Ownership & Lifecycle Management
- Manage and resolve critical P1/P2 incidents and high-priority/high-visibility incidents within SLAs.
- Act as the primary owner and escalation point, ensuring end-to-end lifecycle management (intake, investigation, escalation, resolution, and closure).
- Drive incident prioritization, impact assessment, and correct severity classification.
- Lead coordination and execution of production patches/workarounds required for critical incident resolution.
Incident Coordination & Bridge Management
- Assemble and lead an Incident Response Team (IRT) across engineering, infrastructure, cloud, and support teams.
- Initiate and manage bridge calls, ensuring structured coordination, ownership clarity, and timely follow-ups.
- Drive resolution momentum by actively engaging teams, removing blockers, and ensuring accountability.
Stakeholder & Customer Communication
- Provide timely, structured communication to customers and internal stakeholders covering impact, ETA, progress, and next steps.
- Ensure communication is clear, consistent, and aligned across all channels during major incidents.
- Manage communication for high-visibility issues impacting key customers or business-critical flows.
Status Page Communication (Customer Transparency)
- Own and ensure high-quality Status Page updates for all P1/P2 incidents.
- Provide clear impact descriptions, affected scope, and accurate status progression (Investigating, Identified, Mitigating, Resolved).
- Maintain consistent communication cadence without update gaps.
- Ensure alignment between Status Page, customer emails, and bridge communication.
- Publish explicit resolution updates including service restoration confirmation and summary.
Incident Documentation & Governance
- Ensure accurate and complete incident documentation across Zendesk, Jira, and RCA repositories.
- Maintain audit-ready records with proper traceability and ownership.
- Ensure all incident data is updated in tracking systems and dashboards.
RCA Ownership & Problem Management Alignment
- Drive Root Cause Analysis creation, review, and timely delivery.
- Ensure RCA quality, completeness, and linkage to preventive and corrective actions.
- Track recurring issues and ensure transition into problem management.
- Follow up on corrective actions and backlog items.
Monitoring & Proactive Incident Prevention
- Ensure proactive monitoring of applications and systems to detect issues early.
- Validate alerts and escalate confirmed impact based on business criticality.
- Prevent potential escalations into P1/P2 incidents wherever possible.
Trend Analysis & Continuous Improvement
- Monitor and analyze incident trends and recurring patterns.
- Identify product stability gaps and operational inefficiencies.
- Drive process improvements and support automation initiatives.
Cross-Team Collaboration
- Collaborate with IT, Engineering, Cloud, DevOps, and Infrastructure teams to improve resolution efficiency and system reliability.
- Ensure correct ownership, escalation path, and accountability across teams.
- Coordinate effectively with external vendors and partner teams when required.
Automation & Tooling Improvement
- Support or drive improvements in monitoring, alerting, and ticketing systems.
- Enhance reporting, dashboards, and automation capabilities.
- Reduce manual effort and improve operational efficiency.
Operational Governance, Reporting & Stakeholder Management
- Lead Daily Operations Reviews (DOR) or similar forums with stakeholders, ensuring clear updates on ongoing incidents and operational health.
- Highlight critical risks, blockers, and high-impact issues, ensuring proper prioritization and leadership visibility.
- Provide structured management reporting, including incident summaries, RCA insights, trend analysis, and key metrics.
- Ensure data-driven decision making through accurate reporting and dashboard utilization.
Operational Excellence & Accountability
- Maintain continuous engagement during incidents to ensure timely resolution.
- Drive closure accountability and ensure no gaps in execution.
- Operate in a 24x7 rotational/on-call model to provide uninterrupted incident coverage.
📌 Senior Analyst - Incident Management (Pune)
🏢 Accelya
📍 Pune