07 Oct
|
HCLTech
|
Gautam Buddh Nagar
07 Oct
HCLTech
Gautam Buddh Nagar
Sr Engineer (Tools & Automation)
Gautam Buddha Nagar, Uttar Pradesh
Job Summary
Key Responsibilities
1. 24x7 Monitoring & Operational Coverage
- Provide and sustain round-the-clock monitoring and alert coverage across global business units and regional platforms.
- Ensure seamless operational support as new geographies, applications, and platforms are onboarded.
- Maintain service continuity without coverage gaps during shift transitions and operational escalations.
2. Business-Correlated Monitoring & Early Detection
- Drive the correlation of IT telemetry, infrastructure events, and application alerts with business transaction impacts.
- Enable earlier detection of customer and business-facing issues through proactive monitoring strategies.
3. Major Incident Management (MIM)
- Effectively support Major Incident bridges while maintaining adequate operational coverage.
- Coordinate concurrent incidents without compromising monitoring effectiveness.
- Ensure timely escalation, stakeholder communication, and resolution tracking during critical incidents.
4. Monitoring Optimization & Alert Management
- Continuously review, tune, and optimize alert configurations in partnership with application owners and sustainment teams.
- Identify and reduce false-positive alerts to improve operational efficiency and focus on actionable events.
- Contribute to monitoring maturity initiatives and alert rationalization efforts.
5. Alert Governance & Data Quality
- Ensure accurate and consistent tagging and classification of alerts.
- Maintain monitoring data integrity during periods of high alert volume and operational noise.
- Support reporting accuracy through disciplined alert management practices.
6. Root Cause Analysis (RCA) Ownership
- Drive Root Cause Analysis activities through to final closure.
- Reduce repeated stakeholder engagements by ensuring comprehensive and accurate problem investigation.
7. Incident Ownership & Closure Management
- Maintain clear visibility of incident ownership, status, and closure progress.
- Ensure incidents are actively tracked through resolution rather than limited to notification and escalation activities.
- Drive accountability across support teams for timely closure of operational issues.
8. Stakeholder Coordination & Follow-Through
- Proactively engage sustainment and resolver groups to obtain acknowledgments, updates, and issue resolution.
- Eliminate the need for repeated manual follow-ups through disciplined operational governance.
- Act as a central coordination point during critical operational events.
9. Incident Reporting & Communication
- Produce consistent, accurate, and timely incident reports for operational and leadership stakeholders.
- Standardize reporting formats, communication cadence, and escalation updates across teams.
- Ensure transparency and visibility of operational health and incident status.
10. Severity Assessment & Escalation Management
- Apply established severity criteria consistently across all operational events.
- Make informed decisions regarding escalations, incident creation, and Major Incident declaration.
- Reduce delays caused by uncertainty in impact assessment and incident classification.
11. Shift Handover & Operational Continuity
- Ensure structured and complete handovers between shifts.
- Communicate monitoring concerns, active incidents, known issues, and product updates effectively.
- Maintain continuity of operational ownership across regions and support teams.
12. Business Impact Assessment
- Rapidly assess and communicate business impact during operational incidents.
- Provide leadership with clear understanding of customer, revenue, and operational risks.
- Support data-driven prioritization and decision-making during critical events.
Required Skills & Experience
- Experience in NOC, Command Center, Service Operations, Incident Management, or Monitoring Operations.
- Robust understanding of enterprise monitoring and observability platforms.
- Hands-on experience managing Major Incidents and stakeholder communications.
Key Responsibilities
- Knowledge of ITIL Incident, Problem, and Major Incident Management processes.
- Strong analytical and root-cause investigation skills.
- Excellent communication and executive reporting capabilities.
- Ability to work effectively in a fast-paced 24x7 support environment.
- Experience coordinating across multiple support partners, application teams, and infrastructure teams.
Preferred Qualifications
- ITIL Foundation or equivalent certification.
- Experience with enterprise monitoring platforms, ServiceNow Event Management, or similar tools.
- Experience supporting global business-critical applications and services.
- Familiarity with business service monitoring, observability, and operational governance models.
Success Measures
- Monitoring coverage adherence and operational continuity.
- Reduction in false-positive alerts.
- Improved alert-to-incident correlation accuracy.
- Identifying & proposing Major Incident contributors, before an impact is caused.
- Timely Major Incident management and stakeholder communications.
- RCA completion and closure effectiveness.
- Incident ownership and closure compliance.
- Consistency of reporting and handover quality.
- Improved proactive detection and business-impact visibility.
Skill Requirements
- Knowledge of ITIL Incident, Problem, and Major Incident Management processes.
- Strong analytical and root-cause investigation skills.
- Excellent communication and executive reporting capabilities.
- Ability to work effectively in a fast-paced 24x7 support environment.
- Experience coordinating across multiple support partners, application teams, and infrastructure teams.
Preferred Qualifications
- ITIL Foundation or equivalent certification.
- Experience with enterprise monitoring platforms, ServiceNow Event Management, or similar tools.
- Experience supporting global business-critical applications and services.
- Familiarity with business service monitoring, observability, and operational governance models.
Success Measures
- Monitoring coverage adherence and operational continuity.
- Reduction in false-positive alerts.
- Improved alert-to-incident correlation accuracy.
- Identifying & proposing Major Incident contributors, before an impact is caused.
- Timely Major Incident management and stakeholder communications.
- RCA completion and closure effectiveness.
- Incident ownership and closure compliance.
- Consistency of reporting and handover quality.
- Improved proactive detection and business-impact visibility.
Other Requirements
📌 Sr Engineer (Tools & Automation) (Gautam Buddh Nagar)
🏢 HCLTech
📍 Gautam Buddh Nagar