24 Sep
|
HCLTech
|
Chennai
Administrator (Tools & Automation)
Chennai, Tamil Nadu
Job Summary
Job Responsibilities : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records. : Incident Detection & Front-Line Triage\r\n• Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage.\r\n• Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow.\r\n• Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption.\r\n• Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories.\r\n• Process Activation:
Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds.\r\n• War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams.\r\nService Restoration & Remediation\r\n• Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix.\r\n• Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident.\r\nCommunication, Collaboration & Post-Incident Actions\r\n• Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals.\r\n• Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs.\r\n• L3 E
Key Responsibilities
NA
Skill Requirements
Skill Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized,
jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
Other Requirements
Other Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
📌 Administrator (Tools & Automation) (Chennai)
🏢 HCLTech
📍 Chennai