27 Sep
|
Myntra
|
Bengaluru
Build the Platform Behind AI at Myntra
nMyntra is looking for a Technical Lead – ML Platform & MLOps to help build the next generation of infrastructure that powers machine learning across Myntra.
n
nYou will have the opportunity to design and build foundational ML platform capabilities used across multiple ML teams, influence platform architecture, solve large-scale infrastructure challenges, and establish engineering standards for how ML workloads are built and operated at Myntra.
nWe are looking for a strong hands-on engineer who enjoys building platforms, debugging complex distributed systems, and simplifying infrastructure for hundreds of ML and engineering users.
nWhat You Will Own
nBuild Myntra's ML Platform
n
n
- Architect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.
n
- Create self-service infrastructure that allows Data Scientists and ML Engineers to train, experiment, deploy and operate models without managing underlying infrastructure.
n
- Design reusable platform abstractions, SDKs, APIs and tooling that dramatically improve ML developer productivity.
n
- Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows.
n
- Drive architecture and technology choices for Myntra's ML infrastructure.
n
nDistributed Training & GPU Infrastructure
n
n
- Build infrastructure for large-scale distributed training across CPU and GPU clusters.
n
- Design and operate multi-node and multi-GPU training environments.
n
- Work with distributed computing frameworks such as Ray or equivalent technologies.
n
- Build intelligent scheduling, resource isolation and autoscaling capabilities for ML workloads.
n
- Improve utilization of expensive GPU infrastructure through scheduling, workload optimization and capacity management.
n
- Design fault-tolerant training systems with checkpointing, retries and recovery mechanisms.
n
nML Workflow & Training Platform
n
n
- Build a world-class training platform supporting the complete ML development lifecycle.
n
- Design scalable orchestration for training, feature engineering, validation and deployment workflows.
n
- Build reusable workflow components using technologies such as Airflow, Ray and Kubernetes.
n
- Improve scheduling, dependency management, execution isolation and reliability for thousands of ML workloads.
n
- Enable experimentation across different compute environments without exposing infrastructure complexity to users.
n
nMLOps & Model Lifecycle
n
n
- Build end-to-end ML lifecycle capabilities covering:
n
- Experimentation → Training → Validation → Model Registry → Deployment → Monitoring → Retraining
n
- Build experiment tracking and model management using MLflow or equivalent technologies.
n
- Enable reliable model versioning, approval, rollout and rollback.
n
- Build automated model validation and production-readiness workflows.
n
- Enable reproducible ML workflows across development, staging and production environments.
n
nModel Serving & AI Infrastructure
n
n
- Build highly scalable infrastructure for real-time, batch and asynchronous model inference.
n
- Design model-serving platforms running on Kubernetes and GPU infrastructure.
n
- Optimize serving systems for latency, throughput, availability and cost.
n
- Explore and adopt technologies such as Ray Serve, NVIDIA Triton, vLLM, SGLang or equivalent platforms where appropriate.
n
- Enable production deployment of traditional ML, deep-learning and emerging AI/LLM workloads.
n
nPlatform Reliability & Observability
n
n
- Treat ML infrastructure as a production-grade distributed platform.
n
- Define and drive SLIs, SLOs, availability and reliability standards for ML platform services.
n
- Build deep observability across infrastructure, pipelines, training workloads and inference systems.
n
- Troubleshoot challenging production issues spanning Kubernetes, GPU workloads, distributed systems, networking, storage and ML pipelines.
n
- Drive root-cause analysis and systematically eliminate recurring operational issues.
n
- Design for high availability, fault tolerance and graceful recovery.
n
nCloud, Kubernetes & Infrastructure
n
n
- Design scalable compute, networking and storage infrastructure for ML workloads.
n
- Build and operate ML systems on Kubernetes and cloud platforms.
n
- Automate infrastructure using Terraform or equivalent Infrastructure-as-Code technologies.
n
- Build secure, isolated and reproducible runtime environments.
n
- Drive infrastructure efficiency through autoscaling, workload placement and cost optimization.
n
nML Developer Experience
nA major part of this role is making complex ML infrastructure simple for users.
nYou will:
n
n
- Build developer-facing platforms, SDKs, APIs and abstractions.
n
- Reduce the time required to move an ML experiment into production.
n
- Eliminate repetitive infrastructure work for Data Scientists and ML Engineers.
n
- Build standardized templates and paved roads for ML development.
n
- Improve debugging, discoverability and observability of ML workloads.
n
- Enable teams to focus on models and business problems rather than infrastructure.
n
nTechnical Leadership
nAt E3, we expect you to go beyond implementing individual components.
nYou will:
n
n
- Own architecture and technical direction for major ML Platform initiatives.
n
- Lead complex system-design discussions and technical reviews.
n
- Convert ambiguous problems into scalable platform solutions.
n
- Drive engineering excellence across reliability, scalability, performance and maintainability.
n
- Mentor engineers and raise the technical bar of the team.
n
- Influence architecture across ML, Data, Platform, SRE and Infrastructure teams.
n
- Evaluate emerging technologies and make pragmatic build-vs-buy decisions.
n
- Take critical systems from concept through architecture, implementation and production adoption.
n
nWhat We Are Looking For
nMust Have
n
n
- 6+ years of strong hands-on software/platform engineering experience.
n
- Strong programming skills in Python.
n
- Deep hands-on experience with Kubernetes, containers and Linux.
n
- Experience building or operating large-scale distributed systems or platform infrastructure.
n
- Strong understanding of cloud infrastructure including compute, storage and networking.
n
- Experience building production-grade CI/CD and automation platforms.
n
- Experience with workflow orchestration such as Apache Airflow or equivalent systems.
n
- Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
n
- Strong understanding of ML training and deployment workflows.
n
- Experience with Infrastructure as Code, preferably Terraform.
n
- Robust debugging and production troubleshooting skills.
n
- Experience building systems with monitoring, logging, metrics and alerting.
n
- Strong fundamentals in system design, reliability and distributed computing.
n
nStrong Differentiators
nWe would especially love to meet you if you have worked on:
n
n
- Ray or other distributed computing frameworks
n
- Distributed or multi-node ML training
n
- GPU and multi-GPU infrastructure
n
- Kubernetes-based ML platforms
n
- ML training platforms used by multiple teams
n
- Model serving and inference infrastructure
n
- GPU scheduling and utilization optimization
n
- Large-scale workflow orchestration
n
- ML platform developer experience
n
- Infrastructure cost and performance optimization
n
- Feature platforms or feature stores
n
nGood to Have
n
n
- Databricks, SageMaker, Vertex AI or similar ML platforms
n
- Model monitoring, data drift and automated retraining
n
- NVIDIA Triton, Ray Serve, vLLM or SGLang
n
- LLM training/inference and LLMOps
n
- Vector databases and retrieval infrastructure
n
- Model governance and lineage
n
- OpenTelemetry, Grafana or similar observability ecosystems
n
n
📌 Technical Lead-Machine learning (Bengaluru)
🏢 Myntra
📍 Bengaluru