LEAD DATA ENGINEER - Databricks (Bengaluru)

LEAD DATA ENGINEER - Databricks (Bengaluru)

23 Sep
|
Happiest Minds Technologies
|
Bengaluru

23 Sep

Happiest Minds Technologies

Bengaluru

Job Description: Technical Lead ? Databricks Data Engineering

Role Overview

We are looking for a hands-on Technical Lead with strong experience in Databricks, PySpark, Spark SQL, SQL, and Delta Lake. The Technical Lead will be responsible for solution-design support, development of complex data pipelines, code reviews, technical problem-solving, and mentoring junior developers.

The role requires an individual who can work closely with architects and business stakeholders, translate solution designs into implementable technical components, establish development standards, and ensure the delivery of scalable, reliable, and maintainable data solutions.

The candidate should also have practical experience using Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools to improve development, testing, debugging, documentation, and code quality.

Key Responsibilities

Technical Leadership and Design Support

- Work closely with solution and data architects to understand high-level architecture and translate it into detailed technical designs.
- Support the architect in evaluating solution options, integration patterns, data models, and pipeline-design approaches.
- Prepare and review low-level design documents, data-flow diagrams, source-to-target mappings, interface specifications, and technical implementation plans.
- Make appropriate technical decisions related to pipeline architecture, data processing, performance, scalability, security, and maintainability.
- Identify technical dependencies, risks, assumptions, and constraints during design and development.
- Conduct technical discussions with business analysts, source-system teams, QA teams, DevOps teams, and other engineering stakeholders.
- Estimate development efforts and support sprint planning, task allocation, and delivery tracking.
- Ensure that technical solutions comply with organizational architecture, security, governance, and data-engineering standards.

Hands-on Development

- Design, develop, test, and maintain scalable ETL/ELT data pipelines using Databricks and PySpark.
- Develop complex data-transformation logic using PySpark, Spark SQL, and SQL.
- Build and maintain Delta Lake tables following Bronze, Silver, and Gold layers of the Medallion Architecture.
- Implement Delta Lake capabilities, including:
- MERGE, upsert, and deduplication operations
- Schema enforcement and schema evolution
- Time Travel and version management
- OPTIMIZE and VACUUM
- Incremental and change-based data processing

- Develop reusable PySpark and Python utilities, frameworks, common components, and coding templates.
- Develop and manage Databricks Workflows and Jobs for orchestration, scheduling, dependency management, retries, and alerting.
- Write complex SQL queries for data transformation, reconciliation, validation, analysis, and troubleshooting.
- Integrate Databricks pipelines with external systems using REST APIs.
- Design and support inbound and outbound file integrations using SFTP.
- Process structured and semi-structured data in CSV, JSON, Parquet, and Delta formats.
- Work with AWS services used within the data engineering ecosystem, particularly Amazon S3.
- Implement appropriate exception handling, logging, auditing, monitoring, alerting, and restart/recovery mechanisms.

Code Review and Engineering Quality

- Conduct detailed code reviews for PySpark, Python, Spark SQL, and SQL components developed by the team.
- Ensure code is functionally correct, reusable,



maintainable, secure, and optimized for performance.
- Establish and enforce coding standards, naming conventions, repository structures, branching practices, and documentation guidelines.
- Review pipeline designs and code for partitioning, joins, caching, shuffle operations, data skew, file sizing, and query execution.
- Ensure adequate unit testing, data validation, reconciliation, error handling, and audit controls.
- Identify technical debt and drive code refactoring and continuous engineering improvements.
- Ensure AI-generated code is properly reviewed, tested, secured, and optimized before it is committed or deployed.
- Participate in pull-request reviews and ensure that identified review comments are addressed before code approval.

Team Mentoring and Delivery Support

- Provide technical guidance and day-to-day support to junior and mid-level data engineers.
- Allocate technical tasks based on developer capability, delivery priority, and complexity.
- Mentor junior developers in Databricks, PySpark, Spark SQL, Delta Lake, Python, and data-engineering best practices.
- Conduct technical walkthroughs, knowledge-sharing sessions, pair-programming exercises, and code-quality reviews.
- Help team members troubleshoot complex development, performance, and production issues.
- Review developer estimates and monitor technical progress against sprint and release commitments.
- Identify skill gaps and create structured learning and development plans for junior engineers.
- Promote reusable solutions and consistent engineering practices across the development team.
- Act as the primary technical point of contact for the offshore or delivery team.

Performance Optimization and Production Support

- Troubleshoot pipeline failures, performance bottlenecks, data-quality issues, and production incidents.
- Optimize Spark workloads through effective use of:
- Data partitioning
- Join and broadcast strategies
- Caching and persistence
- Shuffle optimization
- Data-skew handling
- Query execution plans
- Small-file management

- Lead root-cause analysis for critical pipeline and data-quality incidents.
- Review production-readiness checklists and support deployment planning, release validation, and post-production monitoring.
- Ensure that pipelines include appropriate logging, monitoring, alerting, auditing, retry, and recovery mechanisms.
- Prepare and maintain technical documentation, operational procedures, troubleshooting guides, and support runbooks.

Agentic AI?Enabled Development

- Use Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools to improve engineering productivity.
- Apply AI tools for:
- Generating and refactoring PySpark, Python, and SQL code
- Creating unit tests and data-validation scripts
- Debugging pipeline failures and performance problems
- Reviewing code and identifying possible defects
- Generating technical documentation and code explanations
- Accelerating impact analysis and troubleshooting

- Guide junior developers on effective prompting and responsible use of agentic AI development tools.




- Validate AI-generated code for functional correctness, security, data privacy, maintainability, and performance.
- Ensure that sensitive business data, credentials, and proprietary code are handled according to organizational AI and security policies.

Must-Have Skills

- Robust hands-on experience in Databricks-based data engineering.
- Advanced proficiency in PySpark, Spark SQL, and SQL.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Proven experience designing and developing enterprise-scale ETL/ELT data pipelines.
- Strong experience with Delta Lake and Medallion Architecture.
- Hands-on experience with MERGE, schema enforcement, schema evolution, Time Travel, OPTIMIZE, and VACUUM.
- Experience developing and managing Databricks Workflows and Jobs.
- Strong knowledge of Spark performance optimization, including partitioning, joins, caching, shuffle management, data skew, and execution plans.
- Experience preparing or reviewing low-level technical designs and translating architectural designs into development components.
- Strong experience conducting code reviews and enforcing engineering standards.
- Demonstrated experience mentoring junior developers and providing technical leadership to a development team.
- Strong SQL skills, including complex joins, window functions, aggregations, reconciliation, and validation queries.
- Working knowledge of Python for developing reusable utilities, frameworks, and automation scripts.
- Experience integrating data pipelines with REST APIs and SFTP systems.
- Experience working with CSV, JSON, Parquet, and Delta formats.
- Hands-on experience working with Amazon S3.
- Experience implementing logging, exception handling, auditing, monitoring, alerting, and restart/recovery mechanisms.
- Experience troubleshooting and supporting production data pipelines.
- Hands-on knowledge of Cursor, GitHub Copilot, Windsurf, or another agentic AI-enabled development environment.
- Ability to assess AI-generated code and identify incorrect logic, security risks, performance issues, and maintainability concerns.
- Strong technical problem-solving, communication, stakeholder-management, and team-collaboration skills.

Preferred Skills

- Experience with Databricks Unity Catalog, access control, lineage, and data governance.
- Experience with Databricks Asset Bundles, Databricks CLI, or CI/CD-based deployment approaches.
- Basic knowledge of AWS IAM, Secrets Manager, CloudWatch, and Lambda.
- Experience with Git-based development, branching strategies, pull requests, and CI/CD pipelines.
- Knowledge of automated testing frameworks for PySpark and data pipelines.
- Understanding of cloud security, secrets management, and credential-handling practices.
- Experience working in Agile/Scrum delivery environments.
- Databricks or AWS certification would be an advantage.

Education and Experience

- Bachelor?s degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Typically, 8?12 years of overall data engineering experience.
- At least 3?5 years of hands-on experience with Databricks, PySpark, Spark SQL, and Delta Lake.
- Prior experience working as a Technical Lead, Module Lead, or Senior Data Engineer with responsibility for design support, code reviews, and team mentoring.
- Proven experience delivering and supporting enterprise-scale data platforms in production environments.

📌 LEAD DATA ENGINEER - Databricks (Bengaluru)
🏢 Happiest Minds Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer - databricks (bengaluru) / bengaluru