13 Aug
|
ThoughtFocus
|
Kochi
13 Aug
ThoughtFocus
Kochi
Incident Manager – ITIL & ITSM Operations (8-9 Years)
Role Overview
Experienced Incident Manager responsible for driving end-to-end Incident and Major Incident Management across critical enterprise applications, infrastructure, and digital platforms. The role focuses on minimizing service disruptions, improving MTTA/MTTR, enhancing observability, driving automation, and ensuring operational excellence through ITIL-aligned processes. The candidate will lead proactive monitoring, incident response, and continual service improvement initiatives leveraging AIOps and contemporary observability platforms.
Key Responsibilities
Incident & Major Incident Management
- - Lead and govern Incident and Major Incident Management processes across business-critical services.
- Coordinate cross-functional teams during high-priority incidents and service outages.
- Drive rapid incident resolution to achieve MTTA, MTTR, and SLA targets.
- Facilitate incident communication, stakeholder updates, and executive reporting.
- Conduct post-incident reviews, RCA tracking, and improvement actions.
Monitoring, Observability & Event Management
- - Manage enterprise monitoring and observability platforms for proactive issue detection.
- Utilize AppDynamics (AppD), SigNoz, Grafana, Prometheus, Splunk, Application Insights, and SCOM.
- Monitor infrastructure, application, network, and service health to prevent outages.
- Optimize alerting mechanisms and event correlation capabilities.
- Enhance service visibility through dashboards, telemetry, and operational analytics.
Automation, AIOps & Tool Optimization
- - Drive ServiceNow automation including auto-ticketing and incident workflow automation.
- Support CSI Private AIOps platform initiatives for predictive operations.
- Lead monitoring tool consolidation and operational process optimization.
- Improve threshold maturity, alert tuning, and noise reduction strategies.
- Enable automated remediation and event-driven response capabilities.
Problem Management & Continuous Improvement
- - Collaborate with Problem Management teams to eliminate recurring incidents.
- Analyze incident trends and identify systemic service risks.
- Drive continual service improvement initiatives based on operational metrics.
- Improve service reliability, availability, and operational efficiency.
- Establish proactive monitoring and prevention-focused operational practices.
Stakeholder & Operational Governance
- - Collaborate with Application, Infrastructure, Network, and Business teams.
- Lead incident review meetings and service governance discussions.
- Ensure compliance with SLA, OLA, and operational KPIs.
- Provide operational dashboards and performance reporting to leadership.
- Coordinate vendor support teams during critical incidents and escalations.
Required Skills & Qualifications
- - 8-9 years of experience in ITSM Operations, Production Support, or Incident Management.
- Strong expertise in Incident, Major Incident, Problem, and Change Management processes.
- Hands-on experience with ServiceNow ITSM and workflow automation.
- Experience with APM tools including AppDynamics (AppD), SigNoz, or equivalent platforms.
- Strong knowledge of Grafana, Prometheus, Splunk, Application Insights, and SCOM.
- Experience with PagerDuty, event management, and escalation management frameworks.
- Exposure to monitoring tool consolidation, alert optimization, and threshold management.
- Experience supporting AIOps, observability, and proactive operations initiatives.
- Strong analytical, stakeholder management, and communication skills.
- Experience supporting banking, financial services, digital platforms, or large-scale enterprise environments preferred.
Preferred Experience
- - SCOM: Monitoring infrastructure heartbeats and system states across 612+ management packs.
- Application Insights: Monitoring Digital Banking, NuPoint, Identity, and Tokenization platforms.
- Splunk: Operational analytics for Digital Banking, ACH, FedNow, NuPoint Wire, and Imaging platforms.
- PagerDuty: Managing incident escalations across 35+ squads, improving MTTA and MTTR performance.
- OpManager / NetFlow: Monitoring network devices, interfaces, traffic flows, and service availability.
Preferred Certifications
- - ITIL® v4 Foundation / ITIL® Managing Professional
- ServiceNow Certified System Administrator (CSA)
- ServiceNow ITSM Implementation Specialist (Preferred)
- Splunk Core Certified User / Power User
- AppDynamics Associate Certification (Preferred)
- SRE Foundation Certification
- AIOps Foundation Certification
- Microsoft Azure Monitoring / Application Insights Certification
- AWS Cloud Practitioner or Azure Fundamentals (Preferred)
📌 Incident Manager (Kochi)
🏢 ThoughtFocus
📍 Kochi