Architects, deploys, and scales production-grade Generative AI applications and autonomous agents. Bridges AI modeling and distributed systems engineering using cloud-native patterns, microservices, and Kubernetes.
Key Responsibilities
- Model Serving &
- Kubernetes:
Deploy open-weights models (vLLM, TensorRT-LLM) on Kubernetes (EKS/GKE/AKS) with GPU autoscaling and latency optimization.
- Vector &
- RAG Pipelines:
Build high-throughput retrieval architectures using vector databases (Qdrant, Milvus, Pinecone, pgvector) and streaming data pipelines.