This role will lead complex technical programs, establishes scope and milestones, improves cross-team processes, automates reporting, and acts as a technical liaison across stakeholder teams.
Primary Responsibilities
Program Ownership
- Lead one or more major GPU Cluster Health domains such as strategic customer availability, partner repair execution, RMA/spares governance, data/reporting framework, repair workflow improvement, or engineering platform tooling.
- Translate availability gaps into structured programs with scope, owners, milestones, KPIs, risks, dependencies, and executive-ready status.
- Drive weekly operating rhythm for assigned domains, including KPI review, blockers, escalations, decisions, and follow-through.
Customer and Horizontal Leadership
- Support the vertical/horizontal TPM model by owning either a customer vertical or a horizontal functional area.
- For vertical ownership, lead availability and repair strategy for assigned strategic customers and customer clusters.
- For horizontal ownership, lead cross-cutting programs such as NVIDIA/AMD partner tracking, SDE/SRE tooling, RMA/spares feedback loop, SOP tracking, or metrics/reporting framework.
KPI Definition and Governance
- Define KPI measurement methods for repair health, including cluster availability, unavailable-host backlog, repair cycle time, repair success rate, reopen rate, spare availability, RMA loop performance,
partner responsiveness, and SLA/SLO adherence.
- Drive consistent reporting standards across TPMs and support teams.
- Identify trends across multiple workstreams and convert them into prioritized corrective actions.
Cross-Functional Execution
- Partner with India, Morocco, and Mexico SDE/SRE teams to identify tooling and first-level support needs.
- Prioritize tooling requirements that reduce manual repair coordination, improve triage, automate reporting, or accelerate repair handoffs.
- Work with GSL, CHS, CPV, TRS, Warminator, and data center operations to remove repair blockers.
- Support incident-style escalation for clusters at availability risk.
Continuous Improvement
- Lead post-program retrospectives and drive process updates.
- Standardize repeatable playbooks, intake processes, escalation paths, and repair governance artifacts.
- Identify automation opportunities and advocate for prioritization.
Expected Outcomes
- Major repair programs have explicit objectives, metrics, owners, and measurable improvement.
- Leadership has reliable visibility into customer cluster health and repair risk.
- Repair blockers are resolved through structured governance instead of ad hoc escalation.
- Tooling and reporting requirements are translated into actionable SDE/SRE work.
📌 Program Manager 4 (India)
🏢 Oracle
📍 India