Monitor application health, performance, and availability using monitoring tools.
Investigate, troubleshoot, and resolve application-related incidents within agreed SLAs.
Coordinate with development, infrastructure, and business teams for issue resolution.
Maintain knowledge base articles, SOPs, and operational documentation.
Participate in incident, problem, and change management processes.
Provide on-call support and assist during critical production incidents.
Monitor and support server, network, storage, and cloud settings.
Be familiar with any monitoring tool such as Splunk, Dynatrace, Nagios, Datadog, or similar platforms.
Analyze system logs and performance metrics to identify trends and improvement opportunities.
Automate repetitive operational tasks using scripting or automation tools.