04 Oct
|
Questhiring
|
Gurugram
04 Oct
Questhiring
Gurugram
Responsibilities
- Key Responsibilities
- AI Cloud Platform Strategy, Architecture & Leadership
- Define and own the enterprise AI Platform architecture, technical vision, and multi-year roadmap across Azure and GCP.
- Establish reference architectures, engineering standards, design patterns, and governance frameworks for AI, ML, Generative AI, and Agentic AI platforms.
- Partner with business stakeholders and technology leaders to align platform capabilities with strategic business objectives.
- Identify emerging AI technologies and drive innovation through platform modernization and continuous capability enhancement.
- Lead architecture reviews and make technology selection decisions across AI platform ecosystems.
- Drive platform maturity across reliability, scalability, security, governance, developer experience, and operational excellence.
- Hands-on AI Platform Engineering
- Remain actively involved in the design and implementation of critical platform components and complex technical solutions.
- Provide technical leadership for cloud engineering teams through architecture reviews, design workshops, code reviews, and troubleshooting activities.
- Prototype and validate current AI platform capabilities, tooling, and architectural patterns.
- Lead resolution of complex platform, infrastructure, and AI workload challenges.
- Contribute to automation, platform engineering, Infrastructure as Code, and cloud-native engineering initiatives.
- MLOps, LLMOps & Generative AI Platforms
- Define enterprise standards for MLOps and LLMOps capabilities, including model lifecycle management, deployment, monitoring, observability, governance, and operational excellence.
- Architect and implement scalable GenAI platform services supporting GPT, Gemini, Claude, Llama, Mistral, and open-source foundation models.
- Design enterprise Retrieval-Augmented Generation (RAG) frameworks, model evaluation platforms, prompt management solutions,
and guardrail architectures.
- Establish AI observability frameworks covering performance, hallucination detection, model quality, latency, reliability, and cost optimization.
- Drive Responsible AI adoption through governance controls, model validation, compliance, and risk management frameworks.
- Agentic AI & Intelligent Automation
- Define architecture patterns and implementation standards for Agentic AI and multi-agent systems.
- Build and guide implementation of enterprise AI orchestration platforms using LangGraph, CrewAI, AutoGen, Semantic Kernel, MCP, and emerging frameworks.
- Design secure and governed agent deployment models integrating enterprise systems, APIs, business processes, and knowledge repositories.
- Establish human-in-the-loop governance patterns and operational controls for autonomous AI systems.
- Cloud, Kubernetes & Platform Engineering
- Lead architecture and engineering of enterprise Kubernetes platforms on AKS and GKE supporting AI training and inference workloads.
- Design scalable GPU-enabled infrastructure supporting large-scale AI and GenAI workloads.
- Establish best practices for workload isolation, autoscaling, networking, disaster recovery, security, observability, and platform resilience.
- Guide implementation of AI serving technologies including KServe, Ray Serve, vLLM, Triton Inference Server, and distributed inference architectures.
- Drive platform automation, self-service capabilities, and developer productivity improvements.
- DevOps, Automation & Reliability Engineering
- Establish Infrastructure as Code and platform automation standards using Terraform and cloud-native tooling.
- Lead implementation of CI/CD, GitOps, and deployment automation for AI and ML platforms.
- Drive Site Reliability Engineering (SRE) practices, platform observability, operational readiness, and continuous improvement initiatives.
- Define platform KPIs, SLAs, SLOs, and engineering health metrics.
📌 DevOps Engineer (Gurugram)
🏢 Questhiring
📍 Gurugram