Architects, deploys, and scales production-grade Generative AI applications and autonomous agents. Bridges AI modeling and distributed systems engineering using cloud-native patterns, microservices, and Kubernetes.
Key Responsibilities
Model Serving & Kubernetes: Deploy open-weights models (vLLM, TensorRT-LLM) on Kubernetes (EKS/GKE/AKS) with GPU autoscaling and latency optimization.
Vector & RAG Pipelines: Build high-throughput retrieval architectures using vector databases (Qdrant, Milvus, Pinecone, pgvector) and streaming data pipelines.
Agentic Orchestration: Develop multi-agent workflows (LangGraph, CrewAI) and integrate external tool calling protocols (MCP).
LLMOps & Governance: Set up tracing (LangSmith, OpenTelemetry), guardrails (NeMo Guardrails), and FinOps token budget controls.
Core Skill Requirements
Infrastructure: Kubernetes, Docker, Terraform, AWS/GCP/Azure.
Languages & Frameworks: Python (FastAPI), Go or TypeScript, LangChain/LangGraph, LlamaIndex, PyTorch/Hugging Face.
Data & Inference: Vector DBs, vLLM/Ollama, Kafka, PostgreSQL.