24 Sep
|
HCLTech
|
Nagpur
Subject Matter Expert (Tools&Automation;)
Nagpur, Maharashtra
Job Summary
Job Summary : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records. Job Responsibilities : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation:
Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issue
Key Responsibilities
NA
Skill Requirements
Skill Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized,
jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
Other Requirements
Other Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
📌 Subject Matter Expert (Tools&Automation) (Nagpur)
🏢 HCLTech
📍 Nagpur