22 Sep
|
HCLTech
|
Chennai
Administrator (Tools & Automation)
Chennai, Tamil Nadu
Job Summary
Job Responsibilities : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records. Job Description : Incident Detection & Front-Line Triage\\r\\n• Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage.\\r\\n• Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow.\\r\\n• Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption.\\r\\n• Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories.\\r\\n• Process Activation:
Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds.\\r\\n• War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams.\\r\\nService Restoration & Remediation\\r\\n• Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix.\\r\\n• Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident.\\r\\nCommunication, Collaboration & Post-Incident Actions\\r\\n• Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals.\\r\\n• Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs.\\r\\n• L3 E
Key Responsibilities
NA
Skill Requirements
Skill Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized,
jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
Other Requirements
Other Requirement : Incident Detection & Front-Line Triage • Pattern Recognition: Monitor incoming service desk queues to spot sudden spikes in identical user complaints, indicating a potential widespread system outage. • Alert Validation: Review automated infrastructure and application alerts to filter out false positives before triggering the major incident workflow. • Impact Assessment: Interview affected business units quickly to map out the scope, user count, and financial implications of an ongoing disruption. • Ticket Categorization: Ensure all major incident tickets are accurately tagged with the correct high-priority urgency (P1/P2) and functional categories. • Process Activation: Initiate the formal Major Incident Management (MIM) workflow immediately upon identifying breaches of critical service-level thresholds. • War Room Participation: Join live technical crisis bridges to provide real-time updates and collaborate dynamically with cross-functional technical teams. Service Restoration & Remediation • Workaround Deployment: Apply pre-approved temporary workarounds or failover mechanisms to restore business continuity ahead of a permanent fix. • Validation Testing: Conduct comprehensive smoke tests and user acceptance testing to confirm the systems are fully operational before closing the incident. Communication, Collaboration & Post-Incident Actions • Stakeholder Notification: Broadcast standardized, jargon-free status updates to impacted end-users and executive leadership at regular intervals. • Vendor Coordination: Escalate tickets to third-party providers or external software vendors and track their progress against strict contractual SLAs. • L3 Escalation: Package all technical diagnostic data, logs, and troubleshooting steps neatly when escalating unresolved issues to Level 3 engineering teams. • Chronology Tracking: Maintain a meticulous, minute-by-minute timeline of technical actions taken, system behaviors, and milestones throughout the live incident lifecycle. • PIR Contribution: Provide technical root-cause data and timeline logs to the Major Incident Manager for the formal Post-Incident Review (PIR) and Problem Management records.
📌 Administrator (Tools & Automation) (Chennai)
🏢 HCLTech
📍 Chennai