Problem Manager
Senior Associate (AC Sr. Associate) | SDMO Service Delivery & Management Operations .
ROLE SUMMARY
The Problem Manager is responsible for identifying, investigating, and eliminating the root causes of recurring and high-impact infrastructure incidents, reducing business disruption through structured Problem Management, RCA, and continuous service improvement.
You own the end-to-end Problem Management lifecycle — from reactive investigation of major incident outcomes to proactive trend-based problem identification — and drive corrective action across in-scope infrastructure towers in alignment with Client ITSM processes.
KEY RESPONSIBILITIES
Problem Management Lifecycle
- Own and manage the end-to-end Problem Management lifecycle in accordance with ITIL best practices and Client ITSM processes.
- Identify and prioritize recurring and high-impact incidents for problem investigation; maintain a prioritized problem backlog aligned to business impact and risk.
- Classify problems as reactive (post-major-incident) or proactive (trend-driven) and manage each accordingly.
- Ensure problem records are created, documented, and progressed to resolution within agreed timelines and quality standards.
- Validate that corrective and preventive actions (CAPAs) are implemented through the Change Management process and confirmed effective.
Root Cause Analysis (RCA)
- Lead and facilitate structured RCA sessions for recurring or major infrastructure incidents using proven methodologies (5-Why, Fishbone/Ishikawa, Fault Tree Analysis, timeline reconstruction).
- Ensure accurate identification of both technical and process root causes; validate findings with infrastructure tower leads and resolver teams.
- Deliver RCA documentation to agreed quality standards:
timeline, impact analysis, root cause, corrective actions, and preventive measures.
- Collaborate with Incident Management to ensure all closed major incidents are reviewed for problem record creation.
Known Error & Knowledge Management
- Create and maintain Known Error Records (KERs) and workarounds in the ITSM platform; keep the Known Error Database (KEDB) current and accessible.
- Ensure knowledge articles are created and linked to known errors to reduce incident resolution time and support shift-left capability.
- Support Change Advisory Board (CAB) discussions with problem insights and known risk context from open problem investigations.
Governance, Reporting & Continuous Improvement
- Track and report Problem Management KPIs: open problem records, RCA completion rates, CAPA closure status, repeat incident reduction, and known error base growth.
- Produce problem management reports for operational and service governance reviews; identify systemic trends and improvement opportunities.
- Drive proactive problem management by analyzing incident trends, alert patterns, and availability data to identify failure risk before incidents reoccur.
- Recommend process, technical, and monitoring improvements to infrastructure tower leads and the Incident Manager.
- Leverage ServiceNow analytics, Power BI dashboards,
or similar tools to surface and communicate problem trends.
WHAT YOU MUST HAVE
Experience
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 5–8 years of IT operations or managed services experience, with at least 3 years in a Problem Management or operational analysis role.
- Proven track record of leading structured RCA investigations for complex infrastructure failures.
- Experience working across multi-tower infrastructure environments (compute, network, database, EUC, cloud).
- Demonstrated ability to collaborate with Incident, Change, and Release Management to drive CAPAs to closure.
Technical & Process Skills
- Robust working knowledge of ITIL-aligned Problem Management processes and RCA methodologies.
- Hands-on experience with enterprise ITSM platforms, particularly ServiceNow (problem records, KEDB, dashboards).
- Strong analytical skills: ability to interpret incident trend data, identify patterns, and produce governance-ready insights.
- Experience using data analytics or visualization tools (e.g., Power BI, ServiceNow Performance Analytics, Excel) for trend analysis.
- ITIL Foundation certification.
WHAT SETS YOU APART
- ITIL Practitioner (Problem Management module) or equivalent certification.
- Experience with proactive problem management and reliability engineering practices (e.g., SRE principles, failure mode analysis).
- Track record of measurable repeat incident reduction and KEDB growth through structured problem management.
- Familiarity with healthcare IT environments and Epic operational boundaries.
Exposure to automated trend detection, AI-assisted anomaly identification, or predictive problem management
📌 Problem Manager | Senior Associate (Gurugram)
🏢 PwC
📍 Gurugram