Key Responsibilities
- Incident Management (Primary Responsibility).
- Investigate and resolve applicationrelated incidents, including: Application errors.
- Batch / job failures
- Service stop/start issues
- Import/export and archive failures
- Respond to alerts generated via Dynatrace (synthetic and infrastructuretriggered) and validate whether the issue is applicationspecific.
- Execute approved recovery steps such as application/service restarts, configuration validations, and dependency checks.
- Ensure timely updates, proper documentation, and accurate closure of incidents in the ITSM tool.
- Monitoring & Alert Handling
Analyze Dynatrace alerts related to:
- Synthetic monitoring failures
- Application availability degradation
- Backend service dependency failures
- Work with NOC to distinguish application issues vs infrastructure/network issues.
- Reduce repeat incidents by identifying alert patterns and contributing factors.
- Service Requests
- Fulfill applicationrelated service requests, such as:
- Production support requests
- Application configuration changes
- Access validation coordination (with Directory Services)
- Coordinate with business and upstream teams to ensure completeness and SLA adherence.
- Change & Release Support
- Participate in application change windows and vendor releases (e.g., scheduled maintenance, patching, version updates).
- Perform prechange and postchange validation to ensure application health.
- Assist in rollback or recovery actions if postchange issues are detected.
- Support Network Engineering and Infrastructure teams during firewall or storage changes that impact application connectivity.
- Collaboration & Escalation
Work closely with:
- DBA Operations for database job failures and data issues
- Infrastructure/Backup for storage or backup dependencies
- Network Engineering for connectivity or firewallrelated issues
- NOC for initial triage and monitoring alignment
- Escalate compl