21 Aug
|
Myntra
|
Bengaluru
About The Company
Who are we?
Myntra is India’s leading fashion and lifestyle platform, where technology meets creativity. As pioneers in fashion e-commerce, we’ve always believed in disrupting the ordinary.
We thrive on a shared passion for fashion, a drive to innovate to lead, and an environment that empowers each one of us to pave our own way. We’re bold in our thinking, agile in our execution, and collaborative in spirit.
Here, we create MAGIC by inspiring vibrant and joyous self-expression and expanding fashion possibilities for India, while staying true to what we believe in.
We believe in taking bold bets and changing the fashion landscape of India. We are a company that is constantly evolving into newer and better forms and we look for people who are ready to evolve with us.
From our humble beginnings as a customization company in 2007 to being technology and fashion pioneers today, Myntra is going places and we want you to take part in this journey with us.
Working at Myntra is challenging but fun - we are a young and dynamic team, firm believers in meritocracy, believe in equal opportunity, encourage intellectual curiosity and empower our teams with the right tools, space, and opportunities.
Technical Lead –
- ML Platform &
- MLOps (E3)
Location: Bengaluru
Experience: 6+ years
Company: Myntra
Build the Platform Behind AI at Myntra
Myntra is looking for a Technical Lead –
- ML Platform &
- MLOps to help build the next generation of infrastructure that powers machine learning across Myntra.
This is not a conventional MLOps operations role.
You will work on some of the hardest engineering problems at the intersection of Machine Learning, Distributed Systems, Kubernetes, GPU infrastructure, Cloud, Developer Platforms and SRE.
The platform you build will enable Data Scientists and ML Engineers to move from an idea to production faster:
Code →
- Features →
- Distributed Training →
- Experimentation →
- Model Registry →
- Deployment →
- Inference →
- Observability
You will have the prospect to design and build foundational ML platform capabilities used across multiple ML teams, influence platform architecture, solve large-scale infrastructure challenges, and establish engineering standards for how ML workloads are built and operated at Myntra.
We are looking for a strong hands-on engineer who enjoys building platforms, debugging complex distributed systems, and simplifying infrastructure for hundreds of ML and engineering users.
What You Will Own
Build Myntra's ML Platform
- Architect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.
- Create self-service infrastructure that allows Data Scientists and ML Engineers to train, experiment, deploy and operate models without managing underlying infrastructure.
- Design reusable platform abstractions, SDKs, APIs and tooling that dramatically improve ML developer productivity.
- Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows.
- Drive architecture and technology choices for Myntra's ML infrastructure.
Distributed Training &
- GPU Infrastructure
- Build infrastructure for large-scale distributed training across CPU and GPU clusters.
- Design and operate multi-node and multi-GPU training environments.
- Work with distributed computing frameworks such as Ray or equivalent technologies.
- Build intelligent scheduling, resource isolation and autoscaling capabilities for ML workloads.
- Improve utilization of expensive GPU infrastructure through scheduling,
workload optimization and capacity management.
- Design fault-tolerant training systems with checkpointing, retries and recovery mechanisms.
ML Workflow &
- Training Platform
- Build a world-class training platform supporting the complete ML development lifecycle.
- Design scalable orchestration for training, feature engineering, validation and deployment workflows.
- Build reusable workflow components using technologies such as Airflow, Ray and Kubernetes.
- Improve scheduling, dependency management, execution isolation and reliability for thousands of ML workloads.
- Enable experimentation across different compute environments without exposing infrastructure complexity to users.
MLOps &
- Model Lifecycle
- Build end-to-end ML lifecycle capabilities covering: Experimentation →
- Training →
- Validation →
- Model Registry →
- Deployment →
- Monitoring →
- Retraining
- Build experiment tracking and model management using MLflow or equivalent technologies.
- Enable reliable model versioning, approval, rollout and rollback.
- Build automated model validation and production-readiness workflows.
- Enable reproducible ML workflows across development, staging and production environments.
Model Serving &
- AI Infrastructure
- Build highly scalable infrastructure for real-time, batch and asynchronous model inference.
- Design model-serving platforms running on Kubernetes and GPU infrastructure.
- Optimize serving systems for latency, throughput, availability and cost.
- Explore and adopt technologies such as Ray Serve, NVIDIA Triton, vLLM, SGLang or equivalent platforms where appropriate.
- Enable production deployment of traditional ML, deep-learning and emerging AI/LLM workloads.
Platform Reliability &
- Observability
- Treat ML infrastructure as a production-grade distributed platform.
- Define and drive SLIs, SLOs, availability and reliability standards for ML platform services.
- Build deep observability across infrastructure, pipelines, training workloads and inference systems.
- Troubleshoot challenging production issues spanning Kubernetes, GPU workloads, distributed systems, networking, storage and ML pipelines.
- Drive root-cause analysis and systematically eliminate recurring operational issues.
- Design for high availability, fault tolerance and graceful recovery.
Cloud, Kubernetes &
- Infrastructure
- Design scalable compute, networking and storage infrastructure for ML workloads.
- Build and operate ML systems on Kubernetes and cloud platforms.
- Automate infrastructure using Terraform or equivalent Infrastructure-as-Code technologies.
- Build secure, isolated and reproducible runtime environments.
- Drive infrastructure efficiency through autoscaling, workload placement and cost optimization.
ML Developer Experience A major part of this role is making complex ML infrastructure simple for users.
You Will
- Build developer-facing platforms, SDKs, APIs and abstractions.
- Reduce the time required to move an ML experiment into production.
- Eliminate repetitive infrastructure work for Data Scientists and ML Engineers.
- Build standardized templates and paved roads for ML development.
- Improve debugging,
discoverability and observability of ML workloads.
- Enable teams to focus on models and business problems rather than infrastructure.
Technical Leadership At E3, we expect you to go beyond implementing individual components.
You Will
- Own architecture and technical direction for major ML Platform initiatives.
- Lead complex system-design discussions and technical reviews.
- Convert ambiguous problems into scalable platform solutions.
- Drive engineering excellence across reliability, scalability, performance and maintainability.
- Mentor engineers and raise the technical bar of the team.
- Influence architecture across ML, Data, Platform, SRE and Infrastructure teams.
- Evaluate emerging technologies and make pragmatic build-vs-buy decisions.
- Take critical systems from concept through architecture, implementation and production adoption.
What We Are Looking For
Must Have
- 6+ years of strong hands-on software/platform engineering experience.
- Strong programming skills in Python.
- Deep hands-on experience with Kubernetes, containers and Linux.
- Experience building or operating large-scale distributed systems or platform infrastructure.
- Solid understanding of cloud infrastructure including compute, storage and networking.
- Experience building production-grade CI/CD and automation platforms.
- Experience with workflow orchestration such as Apache Airflow or equivalent systems.
- Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
- Strong understanding of ML training and deployment workflows.
- Experience with Infrastructure as Code, preferably Terraform.
- Strong debugging and production troubleshooting skills.
- Experience building systems with monitoring, logging, metrics and alerting.
- Robust fundamentals in system design, reliability and distributed computing.
Strong Differentiators We would especially love to meet you if you have worked on:
- Ray or other distributed computing frameworks
- Distributed or multi-node ML training
- GPU and multi-GPU infrastructure
- Kubernetes-based ML platforms
- ML training platforms used by multiple teams
- Model serving and inference infrastructure
- GPU scheduling and utilization optimization
- Large-scale workflow orchestration
- ML platform developer experience
- Infrastructure cost and performance optimization
- Feature platforms or feature stores
Good to Have
- Databricks, SageMaker, Vertex AI or similar ML platforms
- Model monitoring, data drift and automated retraining
- NVIDIA Triton, Ray Serve, vLLM or SGLang
- LLM training/inference and LLMOps
- Vector databases and retrieval infrastructure
- Model governance and lineage
- OpenTelemetry, Grafana or similar observability ecosystems
The Kind of Engineer Who Will Thrive Here You enjoy going below the abstraction.
You don't just know how to deploy a model using a managed service — you want to understand and improve what happens underneath it.
You Are Equally Comfortable Discussing:
Python →
- Kubernetes →
- Distributed Systems →
- GPUs →
- ML Training →
- CI/CD →
- MLOps →
- Production Reliability
You can take an ambiguous infrastructure problem, design the architecture, build the critical pieces, debug it under production load and create a platform that other engineers love using.
Most importantly, you want to build foundational infrastructure that changes how ML engineering is done at scale.
If building the platform behind Myntra's next generation of AI and ML systems excites you, we would love to talk.
Required Skills machine learning platform, MLP, ML Workflow, data platform
📌 Technical Lead – ML Platform & MLOps (Bengaluru)
🏢 Myntra
📍 Bengaluru