09 Aug
|
EnterOne
|
India
AI Systems & Deployment Engineer
Full-Time • Engineering• AI/ML Infrastructure
ROLE OVERVIEW
We are looking for an AI Systems & Deployment Engineer to own the full lifecycle of large language model deployments — from initial model selection and fine-tuning through production scaling and continuous optimization. You will work at the intersection of cutting-edge ML research and enterprise infrastructure, turning state-of-the-art models into reliable, low-latency services that power real products.
KEY RESPONSIBILITIES
- Deploy and scale open-source and proprietary LLMs using modern inference engines; establish and maintain CI/CD pipelines for ML models in production.
- Fine-tune and optimize models for low-latency inference, including quantization, distillation, and throughput benchmarking.
- Architect advanced Retrieval-Augmented Generation (RAG) pipelines and integrate enterprise vector databases (e.g., Pinecone, Weaviate, pgvector).
- Design and build multi-agent systems — orchestrating tool use, memory, and inter-agent communication for complex autonomous workflows.
- Partner directly with Infrastructure Engineers to ensure GPU cluster efficiency, cost optimization, and uptime SLAs.
- Utilize the NVIDIA software stack (CUDA, TensorRT, Triton Inference Server) to maximize hardware efficiency and throughput.
REQUIRED QUALIFICATIONS
- 3+ years in AI/ML software engineering, with at least 1 year focused on production LLM deployments, fine-tuning, and RAG architectures.
- Hands-on experience with inference frameworks such as vLLM, TGI, TensorRT-LLM, or ONNX Runtime.
- Proficiency in Python and familiarity with MLOps tooling (MLflow, Weights & Biases, DVC, or similar).
- Strong understanding of transformer architectures and common fine-tuning techniques (LoRA, QLoRA, RLHF).
- Experience deploying containerized ML workloads on Kubernetes or equivalent orchestration platforms.
- Ability to evaluate and benchmark models across latency, throughput, and quality dimensions.
PREFERRED QUALIFICATIONS
- Experience with multi-agent orchestration frameworks (LangGraph, CrewAI, AutoGen, or custom implementations).
- Familiarity with NVIDIA Triton Inference Server and GPU profiling tools (Nsight, nvtop).
- Background in distributed systems or high-performance computing.
- Prior work with enterprise vector stores and hybrid search architectures.
WHAT WE OFFER
- Competitive salary and equity in a well-funded, growing company.
- Direct access to cutting-edge GPU infrastructure and the latest foundation models.
- A collaborative team that ships quick and values engineering excellence.
- Flexible remote-friendly work environment.
📌 AI Systems & Deployment Engineer (India)
🏢 EnterOne
📍 India