02 Aug
|
solve IT consultant
|
Chennai
02 Aug
solve IT consultant
Chennai
Job Title: Lead / Principal MLOps Engineer
Job Location: Gurugram or Chennai (5 Days Work From Office)
Experience Level: 8+ Years
Employment Type: Full-time
Department: Artificial Intelligence / Cloud Engineering & Infrastructure
About the Role
We are seeking an experienced and battle-tested MLOps Engineer with 8+ years of expertise to own, scale, and maintain our end-to-end Machine Learning and Generative AI infrastructure. In this role, you will be responsible for ensuring the high availability, reliability, and continuous performance of production ML pipelines, automated monitoring systems, and model deployment workflows.
Working on-site from our Gurugram or Chennai offices, you will bridge the gap between Data Science, Data Engineering, and DevOpsbuilding resilient infrastructure that powers production AI systems at scale.
Key Responsibilities
1. Production ML Infrastructure & Platform Ownership
- Architect, deploy, and manage robust, scalable Machine Learning infrastructure across multi-cloud or hybrid environments (AWS, Azure, or GCP).
- Automate CI/CD pipelines for ML models (CT/CD), ensuring seamless model deployment, rollback strategies, and zero-downtime releases.
- Design and maintain feature stores, model registries, and containerized deployment runtime environments (Docker, Kubernetes/KServe).
1. Pipeline Monitoring & Operational Issue Resolution
- Establish real-time telemetry, monitoring, and alerting frameworks to track system health, inference latency, GPU/CPU utilization, and pipeline throughput.
- Serve as the primary escalation point for production operational incidentsrapidly diagnosing and resolving pipeline failures, data drift, and infrastructure bottlenecks.
- Conduct root-cause analysis (RCA) for operational outages and implement permanent remediations to safeguard system SLAs.
1. System Reliability & Performance Optimization
- Ensure strict adherence to high-availability (99.9%+ uptime), fault tolerance, and disaster recovery standards across all ML workflows.
- Monitor models in production for concept drift, data drift, and latency degradation, triggering automated retraining and re-deployment workflows.
- Optimize resource allocation, cluster auto-scaling, and compute usage to reduce cloud infrastructure costs without compromising speed or reliability.
1. Configuration Management & Security governance
- Implement and manage Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or Ansible to enforce configuration consistency across environments.
- Manage configuration updates, version control, and secret management for complex, distributed ML and LLM microservices.
- Partner with InfoSec teams to enforce data governance, access controls, compliance standards, and security patches across the ML ecosystem.
Requirements & Qualifications
- Experience: 8+ years of qualified engineering experience, with at least 4+ years dedicated to MLOps, ML Platform Infrastructure, or Cloud Engineering at scale.
- Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Core Technical Expertise:
MLOps Frameworks: Hands-on mastery with platforms like MLflow, Kubeflow, Airflow, Weights & Biases, Argo Workflows, or SageMaker. Containerization & Orchestration: Advanced expertise in Docker, Kubernetes (EKS/GKE/AKS), Helm charts, and ingress controllers.
CI/CD & IaC: Strong command of GitHub Actions, GitLab CI, Jenkins, and Infrastructure-as-Code tools like Terraform.
Monitoring & Observability: Proficiency with Prometheus, Grafana, ELK Stack, Datadog, or specialized ML monitoring tools (Evidently AI, Whylogs, Arize).
Programming & Scripting: Expert proficiency in Python, Bash, and SQL for automation, CLI tooling, and service integration.
- Location & Work Mode: Willingness to work 5 days from office at either our Gurugram or Chennai location.
Preferred Qualifications
- Experience with LLMOps (deploying, serving, and monitoring Large Language Models using vLLM, Ollama, or Triton Inference Server).
- Hands-on experience managing GPU compute clusters, CUDA acceleration, and distributed inference/training.
- Relevant certifications in AWS/GCP/Azure Cloud Architecture or Kubernetes (CKA/CKAD).
📌 MLOps Engineer (Chennai)
🏢 solve IT consultant
📍 Chennai