Location - Noida, Bengaluru, Chennai, Hyderabad and Pune.
Role
An AI PRE Engineer (L4) acts as a Platform Reliability Architect and Production Readiness Leader responsible for defining, governing, and scaling enterprise-grade AI/ML platforms. This role ensures AI systems are production-ready, resilient, observable, secure, compliant, and cost-productive at scale, while driving standardization, automation, and platform-wide governance across AI engineering, SRE, DevOps, and MLOps disciplines. The L4 PRE acts as a strategic bridge between platform engineering, AI teams, and business leadership, ensuring reliability and scalability of AI solutions.
Responsibilities
Architecture & Production Readiness Governance
- Define and govern enterprise-wide production readiness frameworks across: Platform, data, model, application, and security layers
- Establish organization-wide PRE standards, policies, and maturity models
- Define and drive release certification governance (PRR frameworks) across all AI workloads
- Own production readiness lifecycle from design → deployment → scale
SLO/SLI Strategy & Reliability Engineering
- Define and govern: SLO/SLI/SLA frameworks for latency, availability, quality, safety, drift
- Error budget strategy at platform and business levels
- Establish reliability engineering models for AI systems
- Align SLOs with: Customer experience, business KPIs, and cost targets
Reference Architecture & Platform Design
- Define and publish: Enterprise AI reference architectures for: LLM applications, RAG pipelines, vector stores, agent frameworks
- Batch, real-time, and streaming inference systems
- Standardize: Platform architecture patterns across multi-cloud and hybrid environments
- Drive platform abstraction and reusable AI services
- Design and architect the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.
Deployment & Release Engineering Governance
- Define enterprise standards for: Safe deployment strategi
📌 AI Platform Engineer (Noida)
🏢 HCLTech
📍 Noida