The Technology Operations Manager is responsible for the day-to-day management and operational reliability of production technology platforms and services. This role ensures the stability, availability, performance, and operational efficiency of infrastructure, cloud platforms, applications, and supporting technology services.
The role manages operational teams responsible for monitoring, incident response, platform operations, and service reliability. It supports the implementation of modern operational practices including automation, observability, and Site Reliability Engineering (SRE) principles.
Working under the direction of the Head of Technology Operations, this role ensures that operational procedures, incident management, service monitoring, and platform maintenance activities are executed effectively to maintain service reliability and meet established SLAs and SLOs.
Key Responsibilities:
Production Operations & Service Reliability
· Manage day-to-day production operations across:
o Cloud platforms (AWS/Azure/GCP)
o SaaS/PaaS platforms
o Core systems and digital platforms
o Integration services and supporting infrastructure
· Ensure operational stability including:
o Service availability
o System performance
o Platform stability
· Monitor and maintain compliance with:
o Service Level Agreements (SLAs)
o Service Level Objectives (SLOs)
o Service reliability targets
· Coordinate operational response for service degradation and outages.
Site Reliability Engineering (SRE) & Modern Ops
· Support the implementation of contemporary operations practices including: