10 Sep
|
Nextloop Technologies
|
India
10 Sep
Nextloop Technologies
India
Gurugram, Chennai | 24/Aug/2026
About the job
We are seeking a highly skilled Lead / Principal MLOps Engineer to design, build, and operate secure, scalable, and reliable enterprise ML platform infrastructure. This role will lead the end-to-end operationalization of machine learning workloads—from training and experimentation through deployment, serving, monitoring, and lifecycle management.
Responsibilties
- Design, build, and maintain scalable, secure, highly available, and reliable MLOps and ML platform infrastructure.
- Lead end-to-end ML pipelines covering training, validation, deployment, serving, monitoring, and lifecycle management.
- Build and manage Kubernetes-based model deployment and serving infrastructure.
- Implement robust CI/CD pipelines for ML applications, models, and platform infrastructure.
- Manage Infrastructure as Code and automate cloud provisioning, configuration, and environment management.
- Design and optimize Kubernetes environments for high availability, scalability, disaster recovery, security, and effective resource utilization.
- Implement monitoring and observability solutions across ML models, pipelines, applications, infrastructure, and platform services.
- Monitor and manage model performance, data drift, concept drift, reliability, and operational health.
- Optimize cloud and GPU infrastructure for performance, scalability, reliability, and cost efficiency.
- Lead troubleshooting, incident management, root cause analysis, and production issue resolution. Collaborate with Data Science, ML Engineering, Platform Engineering, and Cloud teams to improve ML workflows.
- Establish MLOps best practices, engineering standards, architecture patterns,
governance controls, and documentation.
- Provide technical leadership, mentoring, architecture guidance, and infrastructure/code reviews..
Required Qualification
- 8+ years of overall engineering experience, including at least 4 years of hands-on experience in MLOps, ML Platform Engineering, or Cloud Engineering.
- Strong hands-on experience designing and operating production MLOps or ML platform environments.
- Strong understanding of ML lifecycle management: experimentation, model versioning, model registry, training, deployment, serving, monitoring, and governance.
- Hands-on experience with one or more: MLflow, Kubeflow, Apache Airflow, Argo Workflows, Weights & Biases (W&B;), or Amazon SageMaker Advanced experience with Docker and Kubernetes, including EKS, GKE, or AKS.
- Strong practical experience with Helm and KServe for Kubernetes-based application and model deployment/ serving.
- Strong experience with GitHub Actions, GitLab CI, or Jenkins to automate build, test, deployment, and release processes.
- Strong cloud-platform experience in AWS, Azure, or GCP; multi-cloud or hybrid-cloud experience is preferred.
- Strong experience with ML pipeline orchestration, model deployment and serving, model monitoring, data drift and concept drift detection, autoscaling, HA, DR, incident management, and RCA.
- Experience optimizing cloud and GPU infrastructure for performance, scalability, reliability, and cost efficiency.
- Strong problem-solving, technical leadership, communication, stakeholder-management, and mentoring abilities.
Skills
MLOps
Kubernetes
Terraform
Python
MLflow
Gurugram, Chennai, On-site | Contract
📌 Lead / Principal MLOps Engineer (India)
🏢 Nextloop Technologies
📍 India