Lead / Principal MLOps Engineer (India)

Lead / Principal MLOps Engineer (India)

06 Aug
|
solve IT consultant
|
India

06 Aug

solve IT consultant

India

Job Title: Lead / Principal MLOps Engineer

Job Location: Remote (Client location Gurugram/Chennai)

Experience Level: 8+ Years

Employment Type: Full-time Department: Artificial Intelligence / Cloud Engineering Infrastructure About the Role

We are seeking an experienced and battle-tested MLOps Engineer with 8+ years of expertise to own, scale, and maintain our end-to-end Machine Learning and Generative AI infrastructure. In this role, you will be responsible for ensuring the high availability, reliability, and continuous performance of production ML pipelines, automated monitoring systems, and model deployment

Workflows. You will bridge the gap between Data Science, Data Engineering, and DevOps—building resilient infrastructure that powers production AI systems at scale. Key Responsibilities

- Production ML Infrastructure & Platform Ownership

● Architect, deploy, and manage robust, scalable Machine Learning infrastructure across multi-cloud or hybrid environments (AWS, Azure, or GCP).

● Automate CI/CD pipelines for ML models (CT/CD), ensuring seamless model deployment, rollback strategies, and zero-downtime releases.

● Design and maintain feature stores, model registries, and containerized deployment runtime environments (Docker, Kubernetes/KServe).

2.

Pipeline

Monitoring & Operational Issue Resolution

● Establish real-time telemetry, monitoring, and alerting frameworks to track system health, inference latency, GPU/CPU utilization, and pipeline throughput.

● Serve as the primary escalation point for production operational incidents—rapidly diagnosing and resolving pipeline failures, data drift, and infrastructure bottlenecks.

● Conduct root-cause analysis (RCA) for operational outages and implement permanent remediations to safeguard system SLAs.

3.

System Reliability Performance

Optimization

● Ensure strict adherence to high-availability (99.9%+ uptime), fault tolerance, and disaster recovery standards across all ML workflows. ● Monitor models in production for concept drift, data drift, and latency degradation,





triggering automated retraining and re-deployment workflows.

● Optimize resource allocation, cluster auto-scaling, and compute usage to reduce cloud infrastructure costs without compromising speed or reliability.

4.

Configuration

Management & Security governance

● Implement and manage Infrastructure-as-Code (IaC) using Terraform, CloudFormation,

or Ansible to enforce configuration consistency across environments.

● Manage configuration updates, version control, and secret management for complex,

distributed ML and LLM microservices.

● Partner with InfoSec teams to enforce data governance, access controls, compliance standards, and security patches across the ML ecosystem.

Requirements & Qualifications

● Experience: 8+ years of professional engineering experience, with at least 4+ years dedicated to MLOps, ML Platform Infrastructure, or Cloud Engineering at scale.

● Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering,

Information Technology, or a related field. ● Core Technical Expertise:

○ MLOps Frameworks: Hands-on mastery with platforms like MLflow, Kubeflow,

Airflow, Weights & Biases, Argo Workflows, or SageMaker.

○ Containerization & Orchestration: Advanced expertise in Docker, Kubernetes

(EKS/GKE/AKS), Helm charts, and ingress controllers.

○ CI/CD & IaC: Solid command of GitHub Actions, GitLab CI, Jenkins, and

Infrastructure-as-Code tools like Terraform.

○ Monitoring & Observability: Proficiency with Prometheus, Grafana, ELK Stack,

Datadog, or specialized ML monitoring tools (Evidently AI, Whylogs, Arize).

○ Programming & Scripting: Expert proficiency in Python, Bash, and SQL for automation, CLI tooling, and service integration.

Preferred

Qualifications

● Experience with LLMOps (deploying, serving, and monitoring Large Language Models using vLLM, Ollama, or Triton Inference Server).

● Hands-on experience managing GPU compute clusters, CUDA acceleration, and distributed inference/training.

● Relevant certifications in AWS/GCP/Azure Cloud Architecture or Kubernetes

(CKA/CKAD).

📌 Lead / Principal MLOps Engineer (India)
🏢 solve IT consultant
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead / principal mlops engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: lead / principal mlops engineer (india) / india