- Monitor application health, performance, and availability using monitoring tools.
- Investigate, troubleshoot, and resolve application-related incidents within agreed SLAs.
- Coordinate with development, infrastructure, and business teams for issue resolution.
- Maintain knowledge base articles, SOPs, and operational documentation.
- Participate in incident, problem, and change management processes.
- Provide on-call support and assist during critical production incidents.
- Monitor and support server, network, storage, and cloud environments.
- Be familiar with any monitoring tool such as Splunk, Dynatrace, Nagios, Datadog, or similar platforms.
- Analyze system logs and performance metrics to identify trends and improvement opportunities.
- Automate repetitive operational tasks using scripting or automation tools.