We are seeking an experienced **Lead Data Engineer** to join the Marsh Innovation and Data Office (IDO) - Data Strategy team in Mumbai. This hybrid role requires at least two days per week in the office and focuses on designing, building, and maintaining high-performance data pipelines across our Databricks analytics platform.
IDO Data Strategy is a global team of data engineering professionals working to deliver scalable, reliable data pipelines that power business decision-making across Marsh globally. The group has a footprint within the India Knowledge Services (IKS) function at Marsh Global Capability Centre in Mumbai, India. IKS is a hub of several hundred professionals that work with Marsh businesses across a wide range of capabilities.
We will count on you to:
Pipeline Design & Development (Primary Focus – 60%)
Design, build, and maintain production data pipelines in Databricks that transform raw data from ingestion workspace through transformation workspace and into consumption workspace following a lakehouse architecture.
Write clean, optimized SQL and Python code to implement efficient data transformations, ensuring high performance and maintainability across all three workspaces.
Develop and enforce DBT (Data Build Tool) workflows for reproducible, version-controlled data transformations; define and manage data models, tests, and documentation.
Implement CI/CD pipelines using GitHub Actions (or similar) to automate testing, validation, and deployment of data pipeline code across development, staging, and production environments.
Manage code repositories using GitHub, including pull request workflows, code review standards, branch strategies, and version control best practices.
Create comprehensive data mappings, transformation logic, and business rules documentation to ensure reliability and transparency of data lineage across ingestion, transformation, and consumption layers.
Validate data quality and correctness through automated testing frameworks; reconcile datasets and troubleshoot pipeline failures in production environments.
Mentoring & Technical Leadership (Secondary Focus – 40%)
Mentor and train team members on Databricks best practices, Python and SQL optimization techniques, DBT development patterns, and GitHub workflows.
Establish and document coding standards, pipeline architecture patterns, and deployment best practices to promote consistency and quality across the data engineering team.
Lead technical design reviews and architecture discussions for current pipeline initiatives; provide guidance on scalability, performance, and maintainability.
Conduct code reviews for pull requests, ensuring adherence to standards and providing constructive feedback to elevate team capability.
Foster a culture of continuous improvement by sharing lessons learned, emerging technologies, and industry best practices relevant to Databricks and data engineering.
Data Quality & Cross-Workspace Integration (Balanced – Supporting Responsibilities)
Champion data quality and integrity across all three workspaces by defining quality metrics, rules, and monitoring processes.
Collaborate with IT and business teams to ensure seamless data integration between source systems, Databricks platform, and downstream reporting systems.
Identify and resolve data anomalies, validate transformations against business logic, and maintain SLAs for pipeline freshness and accuracy.
Participate in requirements discussions with business stakeholders to translate high-level data needs into technical pipeline specifications (10-20% of time).
Support system improvements by analyzing pipeline performance, identifying bottlenecks, and recommending optimizations to enhance data delivery.
What You Need to Have: -
Education & Experience
Bachelor’s degree in computer science, Information Systems, Engineering, or related field (or equivalent professional experience)
At least 5 years of hands-on experience in data engineering, data pipeline development, or related roles
Demonstrated experience designing and maintaining production data pipelines at scale
Technical Skills (Essential)
Advanced proficiency in SQL – writing complex queries, window functions, CTEs, query optimization, and performance tuning
Strong Python programming skills – writing production-quality, well-tested code for data transformation and orchestration
Hands-on experience with Databricks – workspace configuration, lakehouse design, Delta Lake, Spark SQL, and cluster management
Experience with DBT (Data Build Tool) – data modeling, macro development, testing frameworks, and documentation
Proficiency with GitHub for version control – branching strategies, pull requests, merge workflows, and collaborative development
Solid understanding of CI/CD concepts and experience deploying pipelines through automated testing and release pipelines
Working knowledge of cloud platforms (preferably AWS: S3, EC2, Lambda, CloudFormation, or similar IaC tools)
Familiarity with ETL/ELT design patterns and data architecture concepts (data warehousing, data lakes, lakehouse)
Technical Skills (Advantageous)
Experience with orchestration tools such as Apache Airflow, Prefect, or Databricks Workflows
Proficiency in additional programming languages (Scala, R) or tools (Spark, Kafka)
Experience with data quality frameworks (Great Expectations, dbt tests, custom validation logic)
Knowledge of containerization and infrastructure-as-code (Docker, Terraform)
Experience with RDBMS systems (Oracle, SQL Server, Teradata) and data integration patterns
Professional Competencies
Strong problem-solving and debugging skills with the ability to troubleshoot complex data pipeline issues
Excellent communication skills – ability to explain technical concepts clearly to both engineers and non-technical stakeholders
Self-directed learner with willingness to stay current with emerging technologies and data engineering best practices
Ability to work independently on assigned pipeline development while also collaborating effectively within a team environment
Familiarity with Agile methodologies (Scrum, Kanban) and tools such as Jira or Azure DevOps
What Makes You Stand Out
Demonstrated track record mentoring junior engineers and building high-performing data engineering teams
Experience working with Informatica suite of tools or other enterprise data integration platforms
Familiarity with commercial insurance business domain and insurance industry data structures
Understanding of data governance principles and experience implementing data ownership and stewardship models
Certifications such as Databricks Lakehouse Engineer, AWS Certified Data Analytics, or equivalent
Prior experience establishing data engineering standards and best practices across an organization
Strong track record delivering complex pipeline projects on time and to quality standards
Why Join Our Team
Work with cutting-edge data engineering technologies (Databricks, DBT, cloud platforms) on real-world business problems
Develop deep expertise across a modern analytics technology stack
Mentor and grow the next generation of data engineers
Professional development opportunities, interesting technical challenges, and supportive leadership
Vibrant and inclusive culture where you can collaborate with talented colleagues to create innovative data solutions
Access to a wide range of career opportunities across a global organization
Competitive benefits and rewards to enhance your well-being
About Marsh
Marsh (NYSE: MRSH) is a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $27 billion and more than 90,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information, visit corporate.marsh.com, or follow us on LinkedIn and X.
About Marsh Risk
Marsh Risk is a business of Marsh (NYSE: MRSH), a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $24 billion and more than 90,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information about Marsh Risk, visit marsh.com, or follow us on LinkedIn and X.
📌 Lead Data Engineer (Mumbai)
🏢 Marsh
📍 Mumbai