20 Sep
|
Northern Trust
|
Pune
20 Sep
Northern Trust
Pune
Job Summary
The Manager, Cloud Container Services is responsible for the operational leadership, governance, reliability, and lifecycle management of the enterprise Azure Central Kubernetes Service (CAKS) platform. The role leads a team of Kubernetes platform engineers and operations specialists responsible for delivering a secure, scalable, highly available, and automated container platform that enables application teams to deploy and operate workloads efficiently across the enterprise.
This position serves as the operational owner of the centralized AKS platform, driving platform adoption, service reliability, operational excellence, security compliance, and executive reporting. The role is accountable for providing leadership with actionable insights on platform performance, adoption, operational risks, capacity utilization, cost optimization, and service health. The role aligns closely with the enterprise Cloud Foundation strategy and centralized, multi-tenant AKS operating model.
Key Responsibilities
Platform Operations Service Ownership
- Own and manage the enterprise Azure Central Kubernetes (CAKS) platform.
- Ensure platform availability, stability, scalability, security, and operational excellence.
- Lead Incident, Problem, Change, Capacity, and Availability Management processes.
- Establish operational standards, support models, runbooks, and platform governance controls.
- Drive Kubernetes platform lifecycle management, including upgrades, patching, and end-of-life planning.
- Ensure adherence to enterprise security, compliance, and governance requirements.
Kubernetes Platform Management
- Oversee centralized multi-tenant AKS environments supporting mission-critical and non-mission-critical workloads.
- Manage cluster operations, node pool lifecycle, networking, and platform integrations.
- Ensure secure workload onboarding and platform standardization across business units.
- Partner with engineering teams to implement scalable and reusable Kubernetes platform capabilities.
- Drive platform modernization initiatives, automation, and self-service capabilities.
Reliability Engineering Operational Excellence
- Define and monitor SLAs, SLOs, and operational KPIs.
- Lead major incident management and root cause analysis activities.
- Drive reliability engineering practices and resilience improvements.
- Implement automation-first operational processes leveraging GitHub and Infrastructure as Code.
- Ensure disaster recovery and business continuity readiness for critical Kubernetes workloads.
- Reduce operational overhead through standardization and automated remediation.
Executive Reporting Operational Metrics
- Develop executive dashboards and leadership reporting focused on:
- Platform Availability Uptime
- Cluster Health Metrics
- MTTR Incident Trends
- SLA/SLO Compliance
- Kubernetes Resource Utilization
- Capacity Planning Forecasting
- Node Pool Utilization
- Security Compliance Posture
- Platform Adoption Metrics
- Cost Optimization FinOps Insights
- Support Backlog Service Request Metrics
- Present operational reviews, risk assessments, and strategic recommendations to senior leadership
Security Governance
- Maintain enterprise Kubernetes governance standards.
- Oversee RBAC, Azure AD integration, Managed Identities, and secrets management.
- Ensure enforcement of Azure Policy, security guardrails, and compliance controls.
- Partner with Cybersecurity teams to ensure secure platform operations.
- Lead audit and regulatory compliance activities related to container platforms
Stakeholder Vendor Management
- Serve as the primary operational contact for application teams consuming AKS services.
- Build relationships with Architecture, Infrastructure, Security, DevOps, and Engineering teams.
- Collaborate with Microsoft and strategic partners to optimize platform capabilities.
- Drive customer experience and service improvement initiatives across the Kubernetes ecosystem.
People Leadership
- Lead, mentor, and develop a high-performing Kubernetes platform operations team.
- Establish operational goals and performance expectations.
- Drive talent development, succession planning, and technical capability growth.
- Promote a culture of automation, accountability, innovation, and continuous improvement.
Qualifications
- Bachelor s degree in Computer Science, Information Technology, or related field (Master s preferred).
- 10+ years of experience in data engineering, cloud engineering, container engineering or platform engineering with 15+ years of total work experience, including Azure and/or AWS environments.
- 5+ years of people management experience managing Platform Engineering, DevOps, Infrastructure, or Operations Teams.
- Proven experience managing enterprise-scale cloud data platforms, specifically:
- Azure Kubernetes Service (AKS).
- Kubernetes, Docker, Helm, ArgoCD.
- DevOps, GitHub, Terraform, Hub and Spoke Architecture,
- RBAC, Managed Identities, Azure Key Vault etc..
- Strong experience leading ITIL service management functions, including Incident, Problem, Change, and Service Request Management.
- Experience building and presenting executive operational dashboards and service reviews.
- Proven track record of driving platform reliability, operational transformation, and automation initiatives.
- Experience managing globally distributed teams and complex enterprise environments.
- Hands-on experience with at least one major cloud provider (AWS or Azure) and Cloud Kubernetes Services.
Skills & Competencies
- Deep understanding of Kubernetes architecture, container orchestration, and cloud-native technologies.
- Expertise in Azure Kubernetes Service (AKS) administration and operations.
- Experience managing enterprise-scale multi-tenant Kubernetes platforms.
- Strong knowledge of container technologies including Docker and container registries.
- Experience with GitOps practices using ArgoCD and modern CI/CD pipelines.
- Proficiency with Infrastructure as Code (Terraform and Bicep).
- Robust understanding of cloud networking, including Hub-Spoke architecture, Private Link, Azure Firewall, ingress, and service networking.
- Experience implementing Kubernetes security controls, RBAC, Managed Identities, Azure Policy, and Secrets Management.
- Expertise in observability, monitoring, logging, and performance management using enterprise monitoring platforms.
- Disaster Recovery and Business Continuity Planning
- Vendor and Managed Service Provider Governance.
Key Success Metrics
- Enterprise CAKS platform consistently achieves availability, performance, and security objectives.
- Increased adoption of centralized AKS services across business units.
- High levels of automation, self-service, and operational maturity.
- Reduced operational overhead and improved platform reliability.
- Leadership receives timely, actionable operational insights and recommendations.
- High-performing teams delivering secure, scalable, and resilient Kubernetes platform services.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Manager - Cloud Container Services (Pune)
🏢 Northern Trust
📍 Pune