Head of Platform Reliability (Bengaluru)

Head of Platform Reliability (Bengaluru)

14 Aug
|
NAVA
|
Bengaluru

14 Aug

NAVA

Bengaluru

About Nava Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Head of Platform Reliability to own the deployment, availability, and operational excellence of our AI infrastructure. This leader will be responsible for ensuring that our GPU clusters, platform services, and infrastructure operate reliably in production, while continuously improving automation, incident response, and platform resilience.

What You'll Do Platform Reliability & Operations Own the day-to-day reliability, availability, and operational health of Nava's AI infrastructure platform.

Lead production operations across GPU clusters, networking, storage, orchestration, and platform services.

Define and drive Service Level Objectives (SLOs), Service Level Agreements (SLAs), and uptime targets.

Ensure production environments meet the highest standards of reliability, scalability, and operational excellence. Deployment & Release Management Own the deployment strategy for software, firmware, infrastructure updates, and platform releases.

Design and implement secure rollout mechanisms including phased deployments, canary releases, blue-green deployments, and rollback strategies.

Ensure production changes are executed with minimal customer impact and operational risk.

Establish deployment readiness reviews and operational change management processes.

Incident Management

Lead the incident response function for production systems.

Establish incident management processes, escalation frameworks, and post-incident reviews.

Drive Root Cause Analysis (RCA) and ensure corrective and preventive actions are implemented.

Build a culture of operational learning and continuous improvement.

Reliability Engineering





Identify recurring operational issues and drive long-term engineering fixes rather than temporary workarounds.

Improve system resilience through automation, observability, monitoring, alerting, and self-healing capabilities.

Partner with engineering teams to eliminate reliability bottlenecks and technical debt.

Drive capacity planning, performance optimization, and operational readiness. Cross-Functional Leadership Work closely with Platform Engineering, GPU Cluster Engineering, Networking, SRE, Product, and Customer Success teams.

Ensure operational requirements are embedded into system design from the earliest stages.

Drive operational excellence across multiple engineering functions.

Team Leadership

Build and lead a high-performing Platform Reliability and Site Reliability Engineering (SRE) organization.

Mentor engineering managers and technical leaders.

Foster a culture of accountability, ownership, operational discipline, and customer-first thinking.

What We're Looking For 10–15 years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Operations, or Cloud Operations.

Proven experience leading large-scale production infrastructure teams.

Deep expertise in operating highly available distributed systems and cloud platforms.

Strong understanding of





Kubernetes and container orchestration

Linux systems administration

Infrastructure automation

Monitoring and observability platforms

Incident management and production operations

High-availability architecture and disaster recovery

Experience managing large-scale production deployments and release management.

Strong knowledge of operational metrics including SLAs, SLOs, SLIs, MTTR, and incident response best practices.

Exceptional leadership, stakeholder management, and communication skills.

Nice to Have Experience operating AI infrastructure, GPU clusters, or large-scale inference platforms.

Familiarity with NVIDIA GPU infrastructure, InfiniBand/RDMA networking, and distributed AI workloads.

Experience with GitOps, Infrastructure as Code (Terraform, Ansible), CI/CD pipelines, and production automation.

Exposure to cloud-native platforms, HPC environments, or hyperscale infrastructure.

Why Join

Nava? Shape the future of AI infrastructure - Lead reliability and operational excellence for one of the world’s most advanced AI platforms.

Scale with purpose - Build and grow production systems that power next-generation AI workloads for global customers.

Collaborate with world-class talent - Work alongside top-tier engineers, researchers, and operators solving some of the hardest infrastructure challenges in AI.

Build and lead at the frontier - Establish the operational foundations of Nava’s global AI platform while growing a high-impact reliability engineering organization.

Skills: drive,platforms,infrastructure,availability,reliability,operations,reliability engineering,operational excellence,automation,management

📌 Head of Platform Reliability (Bengaluru)
🏢 NAVA
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: head of platform reliability (bengaluru) / bengaluru