06 Aug
|
solve IT consultant
|
India
06 Aug
solve IT consultant
India
Job Title: Lead / Principal MLOps Engineer
Job Location: Remote (Client location Gurugram/Chennai)
Experience Level: 8+ Years
Employment Type: Full time
Department: Artificial Intelligence / Cloud Engineering Infrastructure
About the Role
We are seeking an experienced and battle-tested MLOps Engineer with 8+ years of expertise to own, scale, and maintain our end-to-end Machine Learning and Generative AI infrastructure. In this role, you will be responsible for ensuring the high availability, reliability, and continuous performance of production ML pipelines, automated monitoring systems, and model deployment
Workflows.
You will bridge the gap between Data Science, Data Engineering, and DevOps—building resilient infrastructure that powers production AI systems at scale.
Key Responsibilities
1. Production ML Infrastructure & Platform Ownership
● Architect, deploy, and manage robust, scalable Machine Learning infrastructure across multi-cloud or hybrid environments (AWS, Azure, or GCP). ● Automate CI/CD pipelines for ML models (CT/CD), ensuring seamless model deployment, rollback strategies, and zero-downtime releases.
● Design and maintain feature stores, model registries, and containerized deployment runtime environments (Docker, Kubernetes/KServe).
1. Pipeline Monitoring & Operational Issue Resolution
● Establish real-time telemetry, monitoring, and alerting frameworks to track system health, inference latency, GPU/CPU utilization, and pipeline throughput. ● Serve as the primary escalation point for production operational incidents—rapidly diagnosing and resolving pipeline failures, data drift, and infrastructure bottlenecks.
● Conduct root-cause analysis (RCA) for operational outages and implement permanent remediations to safeguard system SLAs.
1. System Reliability Performance Optimization
● Ensure strict adherence to high-availability (99.9%+ uptime), fault tolerance, and disaster recovery standards across all ML workflows. ● Monitor models in production for concept drift, data drift, and latency degradation,
triggering automated retraining and re-deployment workflows.
● Optimize resource allocation, cluster auto-scaling, and compute usage to reduce cloud infrastructure costs without compromising speed or reliability.
1. Configuration Management & Security governance
● Implement and manage Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or Ansible to enforce configuration consistency across environments.
● Manage configuration updates, version control, and secret management for complex,
distributed ML and LLM microservices.
● Partner with InfoSec teams to enforce data governance, access controls, compliance standards, and security patches across the ML ecosystem.
Requirements & Qualifications
● Experience: 8+ years of professional engineering experience, with at least 4+ years dedicated to MLOps, ML Platform Infrastructure, or Cloud Engineering at scale.
● Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering,
Information Technology, or a related field.
● Core Technical Expertise:
○ MLOps Frameworks: Hands-on mastery with platforms like MLflow, Kubeflow,
Airflow, Weights & Biases, Argo Workflows, or SageMaker.
○ Containerization & Orchestration: Advanced expertise in Docker, Kubernetes
(EKS/GKE/AKS), Helm charts, and ingress controllers.
○ CI/CD & IaC: Strong command of GitHub Actions, GitLab CI, Jenkins, and
Infrastructure-as-Code tools like Terraform.
○ Monitoring & Observability: Proficiency with Prometheus, Grafana, ELK Stack,
Datadog, or specialized ML monitoring tools (Evidently AI, Whylogs, Arize).
○ Programming & Scripting: Expert proficiency in Python, Bash, and SQL for automation, CLI tooling, and service integration.
Preferred Qualifications
● Experience with LLMOps (deploying, serving, and monitoring Large Language Models using vLLM, Ollama, or Triton Inference Server).
● Hands-on experience managing GPU compute clusters, CUDA acceleration, and distributed inference/training.
● Relevant certifications in AWS/GCP/Azure Cloud Architecture or Kubernetes
(CKA/CKAD).
📌 Lead / Principal MLOps Engineer (India)
🏢 solve IT consultant
📍 India