- Lead and manage P1/P2 major incidents and ensure timely service restoration.
- Act as the single point of contact during critical incidents and lead bridge calls.
- Coordinate with application, infrastructure, network, database, and cloud teams for incident resolution.
- Monitor application and infrastructure health using tools such as Splunk, Datadog, Dynatrace, Grafana, or AppDynamics.
- Drive Root Cause Analysis (RCA) and ensure preventive actions are implemented.
- Track and report SLA, SLO, and service availability metrics.
- Implement automation and operational improvements using Python or Shell scripting.
- Manage incident queues, escalations, and stakeholder communications.
- Collaborate with DevOps and engineering teams to improve system reliability and reduce recurring incidents.
- Support problem management, change management, and ITIL processes.
- Develop and maintain runbooks, playbooks, and operational documentation.
- Analyze trends in incidents, outages, and alerts to improve platform stability.
- Perform capacity planning, performance monitoring, and reliability reviews.
- Ensure 24x7 production support readiness and participate in on-call rotations when required.
- Drive continuous improvement initiatives to enhance service resilience and customer experience.