- Lead the L1/L2/L3 support team responsible for non-core application stability, availability, EOD/BOD batch operations, DR readiness, and database/application monitoring.
- Ensure 99%+ uptime, timely incident resolution, proactive monitoring, and solid collaboration with Engineering, DevOps, Infra, and Business teams.
- Strengthen platform reliability through preventive fixes, automation, governance, and continuous improvement.
Key Result Areas
Supporting Actions
Application Uptime & Reliability
- Maintain ≥99% uptime for non-core applications
- Daily health checks, proactive alerting
- Zero unplanned outages
Batch (EOD/BOD) & Scheduler Management
- ≥98% successful EOD/BOD jobs
- Publish job status reports
- Reduce batch failures by ≥20% YoY
Incident & Problem Management
- ≥95% incidents resolved within SLA
- MTTR reduction by ≥20%
- Zero repeat incidents through preventive actions
DR / BCP Readiness
- 100% DR drill success
- Errorfree switchover/fallback
- DR documentation updated quarterly
API/Gateway Performance
- API gateway success rate >98%
- Monitor latency, timeouts, circuit breaker events
- Ensure certificate renewal before expiry
Team Leadership & Stakeholder Management
- Daily task allocation & monitoring
- Weekly RCA reviews, training & capability building
- Positive stakeholder satisfaction (≥4.3/5)