We are seeking an experienced Lead Data Engineer to join the Marsh Innovation and Data Office (IDO) - Data Strategy team in Mumbai. This hybrid role requires at least two days per week in the office and focuses on designing, building, and maintaining high-performance data pipelines across our Databricks analytics platform.
IDO Data Strategy is a global team of data engineering professionals working to deliver scalable, reliable data pipelines that power business decision-making across Marsh globally. The group has a footprint within the India Knowledge Services (IKS) function at Marsh Global Capability Centre in Mumbai, India. IKS is a hub of several hundred professionals that work with Marsh businesses across a wide range of capabilities.
We will count on you to:
Pipeline Design & Development (Primary Focus – 60%)
- Design, build, and maintain production data pipelines in Databricks that transform raw data from ingestion workspace through transformation workspace and into consumption workspace following a lakehouse architecture.
- Write clean, optimized SQL and Python code to implement efficient data transformations, ensuring high performance and maintainability across all three workspaces.
- Develop and enforce DBT (Data Build Tool) workflows for reproducible, version-controlled data transformations; define and manage data models, tests, and documentation.
- Implement CI/CD pipelines using GitHub Actions (or similar) to automate testing, validation, and deployment of data pipeline code across development, staging, and production environments.
- Manage code repositories using GitHub, including pull request workflows, code review standards, branch strategies, and version control best practices.
- Create comprehensive data mappings, transformation logic, and business rules documentation to ensure reliability and transparency of data lineage across ingestion, transformation, and consumption layers.
- Validate data quality and correctness through automated testing frameworks; reconcile datasets and troubleshoot pipeline failures in production environments.
Mentoring & Technical Leadership (Secondary Focus – 40%)
- Mentor and train team members on Databricks best practices, Python and SQL optimization techniques, DBT development patterns, and GitHub workflows.
- Establish and document coding standards, pipeline architecture patterns, and deployment best practices to promote consistency and quality across the data engineering team.
- Lead technical design reviews and architecture discussions for new pipeline initiatives; provide guidance on scalability, performance, and maintainability.
- Conduct code reviews for pull requests, ensuring adherence to standards and providing constructive feedback to elevate team capability.
- Foster a culture of continuous improvement by sharing lessons learned, emerging technologies, and industry best practices relevant to Databricks and data engineering.
Data Quality & Cross-Workspace Integration (Balanced – Supporting Responsibilities)
- Champion data quality and integrity across all three workspaces by defining quality metrics, rules, and monitoring processes.
- Collaborate with IT and business teams to ensure seamless data integration between source systems, Databricks platform, and downstream reporting systems.
- Identify and resolve data anomalies, validate transformations against business logic, and maintain SLAs for pipeline freshness and accuracy.
- Participate in requirements discussions with business stakeholders to translate high-level data needs into technical pipeline specifications (10-20% of time).
- Support system improvements by analyzing pipeline performance, identifying bottlenecks, and recommending optimizations to enhance data delivery.
What You Need to Have: - Education & Experience
- Bachelor’s degree in computer science, Information Systems, Engineering, or related field (or equivalent professional experience)
- At least 5 years of hands-on experience in data engineering, data pipeline development, or related roles
- Demonstrated experience designing and maintaining production data pipelines at scale
Technical Skills (Essential)
- Advanced proficiency in SQL – writing complex queries, window functions, CTEs, query optimization, and performance tuning
- Strong Python programming skills – writing production-quality, well-tested code for data transformation and orchestration
- Hands-on experience with Databricks – workspace configuration, lakehouse design, Delta Lake, Spark SQL, and cluster management
- Experience with DBT (Data Build Tool) – data modeling, macro development, testing frameworks, and documentation
- Proficiency with GitHub for version control – branching strategies, pull requests, merge workflows, and collaborative development
- Solid understanding of CI/CD concepts and experience deploying pipelines through automated testing and release pipelines
- Working knowledge of cloud platforms (preferably AWS: S3, EC2, Lambda, CloudFormation, or similar IaC tools)
- Familiarity with ETL/ELT design patterns and data architecture concepts (data warehousing, data lakes, lakehouse)
Technical Skills (Advantageous)
- Experience with orchestration tools such as Apache Airflow, Prefect, or Databricks Workflows
- Proficiency in additional programming languages (Scala, R) or tools (Spark, Kafka)
- Experience with data quality frameworks (Great Expectations, dbt tests, custom validation logic)
- Knowledge of containerization and infrastructure-as-code (Docker, Terraform)
- Experience with RDBMS systems (Oracle, SQL Server, Teradata) and data integration patterns
Professional Competencies
- Strong problem-solving and debugging skills with the ability to troubleshoot complex data pipeline issues
- Excellent communication skills – ability to explain technical concepts clearly to both engineers and non-technical stakeholders
- Self-directed learner with willingness to stay current with emerging technologies and data engineering best practices
- Ability to work independently on assigned pipeline development while also collaborating effectively within a team environment
- Familiarity with Agile methodologies (Scrum, Kanban) and tools such as Jira or Azure DevOps
What Makes You Stand Out
- Demonstrated track record mentoring junior engineers and building high-performing data engineering teams
- Experience working with Informatica suite of tools or other enterprise data integration platforms
- Familiarity with commercial insurance business domain and insurance industry data structures
- Understanding of data governance principles and experience implementing data ownership and stewardship models
- Certifications such as Databricks Lakehouse Engineer, AWS Certified Data Analytics, or equivalent
- Prior experience establishing data engineering standards and best practices across an organization
- Strong track record delivering complex pipeline projects on time and to quality standards
Why Join Our Team
- Work with cutting-edge data engineering technologies (Databricks, DBT, cloud platforms) on real-world business problems
- Develop deep expertise across a modern analytics technology stack
- Mentor and grow the next generation of data engineers
- Professional development opportunities, compelling technical challenges, and supportive leadership
- Vibrant and inclusive culture where you can collaborate with talented colleagues to create innovative data solutions
- Access to a wide range of career opportunities across a global organization
- Competitive benefits and rewards to enhance your well-being
About Marsh Marsh (NYSE: MRSH) is a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $27 billion and more than 90,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information, visit corporate.marsh.com, or follow us on LinkedIn and X.
About Marsh Risk
Marsh Risk is a business of Marsh (NYSE: MRSH), a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $24 billion and more than 90,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information about Marsh Risk, visit marsh.com, or follow us on LinkedIn and X.
📌 Lead Data Engineer (Mumbai)
🏢 Marsh
📍 Mumbai