21 Aug
|
Myntra
|
Bengaluru
About The Company
Who are we?
Myntra is India’s leading fashion and lifestyle platform, where technology meets creativity. As pioneers in fashion e-commerce, we’ve always believed in disrupting the ordinary.
We thrive on a shared passion for fashion, a drive to innovate to lead, and an environment that empowers each one of us to pave our own way. We’re bold in our thinking, agile in our execution, and collaborative in spirit.
Here, we create MAGIC by inspiring vibrant and joyous self-expression and expanding fashion possibilities for India, while staying true to what we believe in.
We believe in taking bold bets and changing the fashion landscape of India. We are a company that is constantly evolving into newer and better forms and we look for people who are ready to evolve with us.
From our humble beginnings as a customization company in 2007 to being technology and fashion pioneers today, Myntra is going places and we want you to take part in this journey with us.
Working at Myntra is challenging but fun - we are a young and dynamic team, firm believers in meritocracy, believe in equal opportunity, encourage intellectual curiosity and empower our teams with the right tools, space, and opportunities.
Technical Lead – Data Engineering
Myntra Data Platform (MDP) | Bengaluru | 6+ Years
About The Role
Myntra Data Platform (MDP) powers the data foundation behind analytics, experimentation, personalization, ML and critical business decisions across Myntra.
We are looking for a hands-on Technical Lead – Data Engineering to build and evolve large-scale data processing, ETL and orchestration platforms, while helping shape the next generation of AI-native and agentic data engineering at Myntra.
This is a senior individual-contributor role requiring deep technical ownership—from architecture and production-quality code to performance optimization, troubleshooting and operational excellence.
The ideal candidate combines strong software engineering fundamentals with deep expertise in Spark, SQL, ETL, workflow orchestration, data modeling and modern lakehouse architectures.
What You'll Do
Large-Scale Batch Processing
- Design, build and optimize large-scale batch processing systems using Apache Spark, PySpark and Spark SQL.
- Build production-grade pipelines handling high-volume transformations, joins and aggregations.
- Apply deep understanding of Spark DAGs, stages, shuffles, partitioning, joins, caching, memory and query plans.
- Diagnose data skew, shuffles, memory pressure, inefficient joins and partitioning issues.
- Optimize execution time, compute utilization and cost.
- Build reusable Spark libraries, frameworks and engineering patterns.
ETL & Data Pipeline Engineering
- Build reliable ETL/ELT pipelines across ingestion, transformation, validation, enrichment and publishing.
- Develop reusable pipeline frameworks and abstractions rather than one-off solutions.
- Design for incremental processing, backfills, retries, idempotency, dependencies and failure recovery.
- Ensure correctness, freshness, scalability and maintainability.
- Own pipelines from architecture → development → deployment → production operations.
Workflow Orchestration
- Design and operate large-scale workflows using Apache Airflow, Dagster or equivalent platforms.
- Orchestrate complex dependencies across Spark,
SQL and data-processing workloads.
- Build reusable operators, components and orchestration patterns.
- Improve scheduling, retries, backfills, dependency management and operational visibility.
- Build self-service capabilities for developing, deploying and operating data pipelines at scale.
Databricks & Lakehouse Engineering
- Build and operate large-scale workloads on Databricks and Spark-based environments.
- Design lakehouse architectures using Delta Lake, Apache Iceberg, Parquet or equivalent technologies.
- Optimize partitioning, file sizing, compaction, metadata, Spark and SQL workloads for performance and cost.
- Build platform capabilities that simplify compute, storage and data consumption.
SQL, dbt, Data Transformation & Data Modeling
- Develop and optimize complex production-grade SQL and Spark SQL transformations.
- Build modular transformation models using dbt or equivalent frameworks.
- Establish standards for SQL quality, testing, documentation and dependency management.
- Troubleshoot expensive queries using execution/query plans.
- Design scalable relational, dimensional and lakehouse data models.
- Apply fact/dimension modeling, star schemas, SCDs and incremental models where appropriate.
- Translate business and analytical requirements into reusable data models across raw, curated and consumption-ready layers.
Agentic Data Engineering
- Build AI-native data engineering capabilities across the end-to-end pipeline lifecycle.
- Design AI-assisted workflows that translate requirements into data models, SQL/Spark transformations and executable pipelines.
- Build agents that understand schemas, metadata, lineage, business semantics and existing pipelines.
- Apply LLMs/agents to SQL and pipeline generation, code review, testing, documentation and optimization.
- Build intelligent Spark/SQL optimization and pipeline RCA using execution plans, job metrics, logs, dependencies and historical performance.
- Automate generation of data-quality checks, tests, documentation and lineage metadata.
- Build natural-language/semantic interfaces for data discovery and secure pipeline creation/modification.
- Integrate agents with Airflow/Dagster, dbt, Databricks, Spark and metadata/catalog platforms.
- Design human-in-the-loop controls, validation, guardrails and evaluations for AI-generated SQL, code and pipelines.
- Build reusable agent frameworks, tools and APIs, not isolated AI demos.
Data Quality & Governance
- Embed data quality into pipelines with automated freshness, completeness, schema and consistency checks.
- Establish standards for metadata, ownership, lineage and discoverability.
- Integrate with Unity Catalog or equivalent governance/catalog platforms.
- Design controls for PII, access management, auditing and sensitive data.
- Ensure AI-generated and human-authored pipelines follow the same quality, security and governance standards.
Reliability & Operational Excellence
- Own production health of critical pipelines and platform components.
- Monitor failures, SLAs, freshness, data quality, performance and cost.
- Lead troubleshooting and root-cause analysis of complex incidents.
- Eliminate recurring operational problems through automation, platform improvements and intelligent tooling.
Technical Leadership
- Remain deeply hands-on while providing technical direction.
- Lead architecture/design reviews and review Spark, Python and SQL implementations.
- Mentor engineers across Spark, ETL, SQL, data modeling and production engineering.
- Build reusable frameworks and engineering standards adopted across teams.
- Drive responsible adoption of AI-assisted engineering.
- Lead complex cross-team initiatives and influence architecture without direct authority.
What We're Looking For
- 6+ years of hands-on Data Engineering / Software Engineering experience building production data systems at scale.
- Deep expertise in Apache Spark / PySpark, Spark SQL and Spark internals.
- Strong programming skills in Python, Java or Scala.
- Advanced SQL, query optimization and performance-debugging skills.
- Strong experience building production-grade ETL/ELT pipelines.
- Experience with Airflow, Dagster or equivalent orchestration platforms.
- Experience with Databricks or comparable Spark-based platforms.
- Experience with dbt or SQLMesh
- Strong understanding of data modeling and lakehouse architecture.
- Experience with Parquet, Delta Lake and/or Apache Iceberg.
- Experience with data quality, metadata, lineage and governance.
- Strong production ownership—from design and implementation through operations and RCA.
- Demonstrated technical leadership through architecture, code reviews and mentoring.
Good to Have
- Experience building LLM-powered developer tools, AI agents or engineering automation.
- Understanding of LLMs, tool/function calling, RAG, structured generation and agentic workflows.
- Experience applying AI to SQL generation, data discovery, pipeline generation, debugging or optimization.
- Experience with metadata catalogs, semantic layers, lineage or knowledge graphs for AI context.
- Experience building evaluations and guardrails for AI-generated SQL/code.
- Experience building multi-tenant/self-service data platforms, frameworks or SDKs.
- Experience with Kubernetes, cloud infrastructure and Unity Catalog or similar governance platforms.
What Makes This Role Different You will build the engineering foundations that enable data engineering at Myntra scale — Spark processing, orchestration, ETL frameworks, data modeling, lakehouse architecture, governance and production reliability.
At the same time, you will help redefine data engineering for an agentic world — where agents understand data and dependencies, accelerate the journey from intent to production pipelines, automate repetitive engineering work, and operate within strong human-controlled guardrails.
We are looking for someone comfortable moving between Spark execution plans, Python, SQL, Airflow/Dagster DAGs, dbt models, metadata and platform architecture — and excited about making these workflows dramatically better with AI.
Build the data platform of today — and help define how data engineering will work tomorrow.
Required Skills data platform, design, apache, spark, kafka, problem solving skills
📌 Technical Lead- Data Platform (Bengaluru)
🏢 Myntra
📍 Bengaluru