Senior MLOps Engineer
Experience: 5 Years
Location: ITPL, Bangalore
Work Mode: Work from Office 5 Days a Week
Role: Senior MLOps Engineer
Industry: AI / Machine Learning / Generative AI
Job Summary
We are looking for a hands-on Senior MLOps Engineer / AI Deployment Engineer with strong expertise in MLOps, LLMOps, ML Engineering, Model Deployment, Model Governance, and Model Observability.
The role involves owning the end-to-end lifecycle of Deep Learning models, LLMs, and SLMs, including deployment, scaling, monitoring, optimization, governance, and retirement across cloud, on-premises, hybrid, and air-gapped environments.
Key Responsibilities
- Design, build, and manage MLOps and LLMOps pipelines.
- Deploy, host, scale, and optimize Deep Learning models, LLMs, and SLMs.
- Manage the complete model lifecycle including versioning, deployment, rollout, rollback, monitoring, and retirement.
- Deploy and manage models on Kubernetes, OpenShift, Databricks, and GPU-based infrastructure.
- Implement Model Governance, Model Registry, lineage, approval workflows, and compliance controls.
- Build and maintain model monitoring, observability, tracing, logging, and drift detection solutions.
- Optimize model latency, throughput, GPU utilization, inference performance, and infrastructure cost.
- Work across Cloud, On-Premises, Hybrid, and Air-Gapped environments.
- Implement secure and scalable AI/ML inference platforms and deployment architectures.
- Collaborate with Data Scientists, ML Engineers, DevOps, Platform Engineering, and Security teams.
Must-Have Skills
- 35 years of hands-on experience in MLOps, LLMOps, ML Engineering, or AI Engineering.
- Strong Python development skills.
- Hands-on experience with Databricks and/or Azure ML.
- Strong understanding of Deep Learning, LLMs, SLMs, RAG, and Hugging Face.
- Experience deploying models developed using PyTorch and TensorFlow.
- Strong hands-on experience with model deployment on:
- Kubernetes
- Databricks
- GPU Infrastructure
- OpenShift
LLM / Model Serving
Experience with one or more of the following:
- vLLM
- Triton Inference Server
- Ray Serve
- SGLang
- Databricks Model Serving
GPU & Inference Optimization
- Strong knowledge of NVIDIA GPUs and CUDA.
- Experience with multi-GPU deployments.
- Hands-on experience with LLM inference optimization, GPU utilization, latency, and throughput optimization.
Model Governance & Observability
- Model Registry and Model Lifecycle Management.
- Model Governance and Model Lineage.
- Model Monitoring and Drift Detection.
- AI/ML Observability and Distributed Tracing.
- Experience with:
- MLflow
- OpenTelemetry
- Langfuse
- Splunk
- Grafana / ELK
Databases & Vector Databases
Strong database knowledge in one or more of:
- SQL Server
- PostgreSQL
- Oracle
- MySQL
- MongoDB
Experience with Vector Databases / Search platforms such as:
- Pinecone
- Chroma
- FAISS
- Milvus
- Azure AI Search
API & Integration
- REST APIs
- WebSockets
- Streaming HTTP
- Experience building scalable AI/ML inference APIs.
DevOps / CI-CD
- Jenkins
- Azure DevOps
- CI/CD automation for ML and AI workloads.
- Containerization and orchestration using Docker and Kubernetes.
Infrastructure & Security
- Experience working across Cloud, On-Premises, Hybrid, and Air-Gapped environments.
- Experience with authentication and authorization solutions such as Keycloak.
- Understanding of Model Security, Governance, and Compliance.
Positive-to-Have Skills
- Kafka, RabbitMQ, Azure Event Hub.
- Fine-tuning and model optimization.
- LLM quantization and inference optimization.
- Model Governance and AI Security.
- Experience with open-source and enterprise LLMs such as:
- Llama
- Mistral
- DeepSeek
- Qwen
- Phi
- Gemma
Warm regards,
Laya Guptha | Senior IT Recruiter
Email:
[email protected]
📌 Senior MLOps Engineer (Bangalore Rural)
🏢 Tranzeal
📍 Bangalore Rural