27 Sep
|
algoleap
|
India
Job Description – MLOps / LLM Infrastructure Engineer
Experience: 7–12 years
Role Overview
Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud workplace. The role focuses on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.
Key Responsibilities
Deploy and operate models such as GPT, LLaMA, Gemma, Mistral, and other product/open-weight models within sovereign cloud.
Manage GPU provisioning, capacity planning, utilization, and performance optimization.
Implement LLM inference serving using platforms such as vLLM, Triton, or similar frameworks.
Build model deployment, versioning, rollback, and lifecycle management processes.
Develop MLOps pipelines for model packaging, testing, deployment, and monitoring.
Monitor latency, throughput, GPU utilization, availability, and inference costs.
Implement scalable and highly available model-serving infrastructure using Kubernetes and containers.
Work closely with platform, security,
and gateway teams to ensure secure model access and governance.
Troubleshoot production issues across GPU, inference, Kubernetes, networking, and model-serving layers.
Required Skills
Robust experience in MLOps / LLM infrastructure / model serving.
Hands-on experience with GPU-based inference and Kubernetes.
Experience with vLLM, NVIDIA Triton, TensorRT-LLM, or equivalent.
Experience deploying and managing LLaMA, Gemma, Mistral, GPT or similar LLMs.
Knowledge of Docker, Kubernetes, CI/CD, model registries, and observability.
Understanding of LLM quantization, batching, caching, GPU memory management, and inference optimization.
Experience operating ML/LLM workloads in private, on-premises, or sovereign-cloud environments.
Preferred: Experience with NVIDIA GPUs, CUDA, Helm, Prometheus/Grafana, MLflow, and automated model deployment pipelines.
📌 Mlops / Llm Infra Engineers Mumbai (India)
🏢 algoleap
📍 India