01 Oct
|
Trianz
|
Bengaluru
ABOUT THIS ROLE
You will own the complete architecture for private and sovereign AI deployment at Trianz. This is not a cloud-API-consumption role. You will design the system that runs large language models inside customer environments -- choosing the right models, designing the serving topology, setting CPU and GPU routing policies, and ensuring the entire stack is secure, sovereign, and provider-agnostic. This is a pure IC role with architectural authority over how AI runs in production.
WHAT YOU WILL DO
- Design the complete private LLM serving architecture: model selection, serving framework (vLLM,TensorRT-LLM, Triton), and runtime topology.
- Define intelligent CPU vs GPU routing policies: which model sizes and prompt types route to which compute tier based on latency, cost, and throughput targets.
- Architect multi-cloud model serving: AWS, Azure, GCP -- provider-agnostic, no managed AI lock-in.
- Design on-premises serving for enterprise customers: RedHat OpenShift, VMware -- containerised model serving on customer hardware.
- Architect per-tenant model isolation, data-residency compliance, and air-gapped sovereign deployment patterns.
- Define the release architecture for model versions: rollout, staged deployment, rollback, and promotion gates.
- Design the LLM governance framework: model behaviour monitoring, inference audit logging, guardrail architecture.
- Design the DevSecOps pipeline architecture and self-service deployment automation standards.
- Evaluate open-source models (Llama, Mistral, Qwen, Phi) against closed models for specific enterprise use cases.
- Produce architecture sign-off documents and review all AI system designs before implementation.
MUST HAVE
- Deployed LLMs to production in a real enterprise workplace -- not just API consumption.
- Hands-on with vLLM, TensorRT-LLM, or Triton in a production serving context
- Designed CPU cluster inference (Intel Xeon / AMD EPYC) for open-source models.
- Kubernetes at production scale (EKS, AKS, GKE, or OpenShift) -- not just local k8s.
- Designed multi-cloud, provider-agnostic AI architectures.
- Experience with model quantization (GPTQ, AWQ, GGUF) and routing trade-offs
- 8+ years in AI/software architecture with at least 3 years in production LLM systems.
GOOD TO HAVE
- Experience with RedHat OpenShift for on-premises model serving.
- Familiarity with NVIDIA Dynamo, SGLang, or custom inference schedulers.
- GPU FinOps -- cost modelling for mixed CPU/GPU inference fleets.
- Sovereign AI or regulated-industry deployment experience.
- Intel OpenVINO or AMD ROCm for CPU-optimised inference.
Company Overview
Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model". With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.
With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps - delivered through strategic partnerships with leading hyperscalers.
We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence - RevolutionAIzing Transformations.
📌 Principal AI Architect (Bengaluru)
🏢 Trianz
📍 Bengaluru