01 Oct
|
Capgemini
|
Bengaluru
01 Oct
Capgemini
Bengaluru
MLOps Engineer to lead the design and implementation of scalable, secure, and production-grade ML platforms across multi-cloud environments (Azure, AWS, GCP).
The role involves architecting end-to-end ML lifecycle solutions, enabling reproducibility, governance, and operational excellence while working closely with data science, platform engineering, and cloud teams.
1. overall experience and dedicated MLOps / AI Platform Engineering experience
2. Kubernetes hands-on expertise
3. Python development experience
4. CI/CD implementation experience
5. Monitoring and observability experienceRole & responsibilities
We are looking for a lead MLOps Engineer lead the design and implementation of scalable, secure, and production-grade ML platforms across Azure platform. The role involves architecting end-to-end ML lifecycle solutions, enabling reproducibility, governance, and operational excellence while working closely with data science, platform engineering, and cloud teams.
Must-Have Skills
1. Cloud Expertise
- Experience in
- Azure: Azure ML, Azure DevOps, AKS, Functions, Logic Apps, Key Vault
- Databricks: Databricks asset bundles, Unity Catalog, Lakehouse Monitoring, Databricks clusters
2. CI/CD for ML and LLMs
- Build and manage ML pipelines using:
- Azure DevOps / Databricks Asset Bundles
- Implement
- Model build, test, validation, and deployment workflows
- Integration with container registries and artifact stores
3. Model Lifecycle Management
- End-to-end lifecycle:
- Experiment tracking (MLflow)
- Model versioning, registry, packaging, deployment, rollback
- Exposure to model governance and lineage tracking Databricks Unity Catalog and Lakehouse monitoring
4. Infrastructure as Code (IaC)
- Hands-on experience with:
- Terraform (preferred cross-cloud standard)
- Automating infrastructure provisioning for ML workloads
5. Containerization & Orchestration
- Strong experience in:
- Docker-based packaging
- Kubernetes platforms: AKS
- Deploying scalable ML inference services
6. Automation & Orchestration
- Build automated workflows for:
- Training, validation, deployment, retraining
- Experience with orchestration tools
- Airflow / Azure Container Apps
7. ML Observability & Monitoring
- Monitoring tools across cloud ecosystems:
- Azure Monitor, Application Insights
- Open-source: Prometheus, Grafana
- Implement
- Model performance, drift detection, alerting
8. Feature Store & Data Integration
- Experience with feature stores:
- Azure / Databricks Feature Store
- Designing reusable, governed feature pipelines
9. Model Deployment Platforms
- Experience deploying models using: Azure ML
- REST API-based inference endpoints and microservices
10. Programming & APIs
- Strong Python skills (automation, pipelines, ML integration)
- Experience building and consuming REST APIs / microservices
11. Scrum + Agile Delivery
- Certified Scrum Master with 2+ years experience facilitating agile ceremonies and managing structured sprint execution,
- Experience in removing blockers and enabling cross-functional team collaboration
- Experience in aligning sprint outcomes with product roadmap and driving continuous improvement
Positive-to-Have Skills
- Multi-cloud MLOps framework/tooling contributions
- Experience with
- Kubeflow, ZenML, KServe, ONNX, Triton Inference Server
- Security and vulnerabilities
- Exposure to
- Real-time/streaming ML (Kafka, Kinesis, Pub/Sub)
- Responsible AI / governance frameworks
- Cost optimization (FinOps practices across clouds)
Preferred candidate profile
📌 Capgemini Looking For Mlops Lead For Bangalore/Hyderabad/Pune (Bengaluru)
🏢 Capgemini
📍 Bengaluru