13 Aug
|
Randstad
|
Hyderabad
13 Aug
Randstad
Hyderabad
Role & responsibilities
Cloud & Data & AI/ML Platform Operations
- Own the operational health, availability, and performance of enterprise cloud platforms (AWS, Azure, GCP) and AI/ML infrastructure ensuring production-grade reliability for models, pipelines, data & analytics platforms, and cloud-native applications.
- Establish and enforce MLOps and AIOps operational standards model monitoring, drift detection, automated retraining pipelines, inference infrastructure management, and incident response for AI/ML workloads.
Site Reliability Engineering (SRE)
- Build, lead, and mature the enterprise SRE function embedding reliability engineering principles (SLOs, SLIs, error budgets, chaos engineering) across critical digital platforms and services.
- Lead the post-incident review (PIR) and blameless retrospective culture, ensuring every significant incident drives lasting systemic improvements rather than short-term fixes.
Managed Service Provider (MSP) Governance
- Serve as the executive owner of all MSP and third-party operational vendor relationships governing contracts, SLAs, performance metrics, and strategic alignment across managed infrastructure, cloud, support, and security services.
- Lead structured QBRs, performance reviews, and executive-level escalations with MSP partners holding providers accountable to contractual commitments while fostering team-oriented, long-term partnerships.
Enterprise IT Support & Service Management
- Oversee the enterprise IT support function Tier 1/2/3 support, service desk operations, and application support ensuring exceptional end-user experience and first-contact resolution metrics.
- Lead the continuous maturation of ITSM processes (Incident, Problem, Change, Release, and Configuration Management) in alignment with ITIL best practices and enterprise risk controls.
Executive Visibility, Reporting & Stakeholder Engagement
People Leadership, Mentoring & Team Culture
Technical Competencies
- Deep expertise in cloud platform operations (AWS, Azure, GCP) architecture patterns, operational tooling, FinOps, and multi-cloud governance.
- Strong grounding in AI/ML operations MLOps pipelines, model monitoring, inference infrastructure, and AIOps platform tooling.
- SRE mastery — observability stacks (Datadog, Dynatrace, Prometheus/Grafana), chaos engineering, SLO frameworks, and incident management platforms.
- ITSM fluency — ServiceNow or equivalent, ITIL v4 processes, and enterprise support operations at scale.
- MSP governance and vendor management — SLA construction, performance metrics, contract lifecycle, and strategic sourcing principles.
- Security and compliance operations awareness — vulnerability management, cloud security posture, and regulatory compliance frameworks.
📌 Senior Director- Infrastructure, Operations & App Support (Hyderabad)
🏢 Randstad
📍 Hyderabad