26 Sep
|
AlgoLeap Technologies
|
Mumbai
26 Sep
AlgoLeap Technologies
Mumbai
SUMMARY
Job Description MLOps / LLM Infrastructure Engineer
Experience: 7 12 years
Role Overview:
Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud setting. The role focuses on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.
Key Responsibilities:
Deploy and operate models such as GPT, LLaMA, Gemma, Mistral, and other product/open-weight models within sovereign cloud.
Manage GPU provisioning, capacity planning, utilization, and performance optimization .
Implement LLM inference serving using platforms such as vLLM, Triton, or similar frameworks .
Build model deployment, versioning, rollback, and lifecycle management processes.
Develop MLOps pipelines for model packaging, testing, deployment, and monitoring.
Monitor latency, throughput, GPU utilization, availability, and inference costs .
Implement scalable and highly available model-serving infrastructure using Kubernetes and containers .
Work closely with platform, security,
and gateway teams to ensure secure model access and governance.
Troubleshoot production issues across GPU, inference, Kubernetes, networking, and model-serving layers .
Required Skills:
Robust experience in MLOps / LLM infrastructure / model serving .
Hands - on experience with GPU-based inference and Kubernetes .
Experience with vLLM, NVIDIA Triton, TensorRT-LLM, or equivalent.
Experience deploying and managing LLaMA, Gemma, Mistral, GPT or similar LLMs .
Knowledge of Docker, Kubernetes, CI/CD, model registries, and observability .
Understanding of LLM quantization, batching, caching, GPU memory management, and inference optimization .
Experience operating ML/LLM workloads in private, on-premises, or sovereign-cloud settings .
Preferred: Experience with NVIDIA GPUs, CUDA, Helm, Prometheus/Grafana, MLflow, and automated model deployment pipelines .
📌 Mlops / Llm Infra Engineers Mumbai
🏢 AlgoLeap Technologies
📍 Mumbai