Skills: Incident Management, Root Cause Analysis, Team Leadership, Service Level Agreements, Operations Management, Problem-solving
Job Summary
Platform Management: Oversee APM (Application Performance Monitoring) and Infrastructure monitoring stacks. Incident & RCA Management: Define alerting thresholds, reduce alert fatigue, and lead Problem Management/Root Cause Analysis. Team Leadership: Mentor engineers and manage global shift rotations for continuous monitoring. Service Level Agreements (SLA): Ensure systems meet uptime targets and compliance standards Operations Management with focus on continuous improvement, problem-solving, meeting client SLAs, and empowering teams through effective people management.
Key Responsibilities
1. Enhance operational systems to facilitate improved management reporting, streamline information flow, optimize business processes, and support organizational planning.
2. Understand client requirements and accountable in ensuring support team is meeting client expectations.
3. To lead and mentor the project team and ensure transparent communication of project goals.
4. Bringing new ideas and innovation for process development and overall organizational progress.
5.
To provide solutions commensurate with the customers’ needs within the ambit of the given setting so as to lead to business results.
Platform Management: Oversee APM (Application Performance Monitoring) and Infrastructure monitoring stacks. Incident & RCA Management: Define alerting thresholds, reduce alert fatigue, and lead Problem Management/Root Cause Analysis. Team Leadership: Mentor engineers and manage global shift rotations for continuous monitoring. Service Level Agreements (SLA): Ensure systems meet uptime targets and compliance standards.
Skill Requirements
Platform Management: Oversee APM (Application Performance Monitoring) and Infrastructure monitoring stacks. Incident & RCA Management: Define alerting thresholds, reduce alert fatigue, and lead Problem Management/Root Cause Analysis. Team Leadership: Mentor engineers and manage global shift rotations for continuous monitoring. Service Level Agreements (SLA): Ensure systems meet uptime targets and compliance standards.
Other Requirements
Platform Management: Oversee APM (Application Performance Monitoring) and Infrastructure monitoring stacks. Incident & RCA Management: Define alerting thresholds, reduce alert fatigue, and lead Problem Management/Root Cause Analysis. Team Leadership: Mentor engineers and manage global shift rotations for continuous monitoring. Service Level Agreements (SLA): Ensure systems meet uptime targets and compliance standards.