Architects, deploys, and scales production-grade Generative AI applications and autonomous agents. Bridges AI modeling and distributed systems engineering using cloud-native patterns, microservices, and Kubernetes.
Key Responsibilities
- Model Serving & Kubernetes: Deploy open-weights models (vLLM, TensorRT-LLM) on Kubernetes (EKS/GKE/AKS) with GPU autoscaling and latency optimization.
- Vector & RAG Pipelines: Build high-throughput retrieval architectures using vector databases (Qdrant, Milvus, Pinecone, pgvector) and streaming data pipelines.
- Agentic Orchestration: Develop multi-agent workflows (LangGraph, CrewAI) and integrate external tool calling protocols (MCP).
- LLMOps & Governance: Set up tracing (LangSmith, OpenTelemetry), guardrails (NeMo Guardrails), and FinOps token budget controls.
Core Skill Requirements
- Infrastructure: Kubernetes, Docker, Terraform, AWS/GCP/Azure.
- Languages & Frameworks: Python (FastAPI), Go or TypeScript, LangChain/LangGraph, LlamaIndex, PyTorch/Hugging Face.
- Data & Inference: Vector DBs, vLLM/Ollama, Kafka, PostgreSQL.