- Basic understanding about database technology and having experience of granting access related to DB servers
- Continuously monitor infrastructure, application and batch processing alerts using Dynatrace, SolarWinds and Control-M and ensure timely acknowledgment of all events.
- Perform first-level incident triage, impact assessment and categorization based on severity, business impact and service criticality.
- Execute documented SOPs, runbooks and recovery procedures to restore services wherever standard operational actions are approved.
- Analyze alerts, logs, dashboards and monitoring trends to identify probable causes and provide meaningful information during escalations.
- Create, update and manage incident tickets while ensuring SLA compliance and accurate documentation of troubleshooting activities.
- Coordinate with L2/L3, application, infrastructure and business teams for resolution of high-priority incidents and service disruptions.
- Monitor Control-M jobs, identify job failures,
perform defined recovery actions and escalate unresolved batch issues.
- Participate in major incident bridges, provide status updates and communicate outage information through established processes.
- Prepare shift handover reports, operational summaries and recurring issue trends to support continuous service improvement.
- Maintain operational documentation and contribute to knowledge management repositories.
Desired Experience:
- 2-5 years in application/infrastructure monitoring support, especially who have experience working in NOC team
- Experience in Windows Server, AD, DNS/DHCP; Dynatrace; SolarWinds; Control-M; ServiceNow
Soft Skills:
- Strong analytical and troubleshooting skills
- Positive communication and stakeholder coordination
- Ability to work in 24/7 shifts and handle production environments