DATA ENGINEER - Databricks (Bengaluru)

DATA ENGINEER - Databricks (Bengaluru)

08 Sep
|
Happiest Minds Technologies
|
Bengaluru

08 Sep

Happiest Minds Technologies

Bengaluru

: Senior Databricks Data Engineer

Role Overview

We are looking for an experienced Senior Databricks Data Engineer to design, develop, and maintain scalable ETL/ELT data pipelines using Databricks, PySpark, Spark SQL, and Delta Lake. The candidate should have strong hands-on experience in data transformation, pipeline orchestration, performance optimization, data quality, and production support within an AWS-based data engineering environment.

The role also requires practical experience using Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools to accelerate development, testing, debugging, documentation, and code optimization.

Key Responsibilities

- Design, develop, test, and maintain scalable ETL/ELT data pipelines using Databricks and PySpark.
- Develop and optimize data transformation logic using PySpark, Spark SQL, and SQL.
- Build and maintain Delta Lake tables following Bronze, Silver, and Gold layers of the Medallion Architecture.
- Implement Delta Lake capabilities, including:
- MERGE and upsert operations
- Schema enforcement and schema evolution
- Time Travel and version management
- OPTIMIZE and VACUUM
- Incremental and change-based data processing

- Develop and manage Databricks Workflows and Jobs for pipeline orchestration, scheduling, dependency management, retries, and alerting.
- Design productive incremental-load and restart/recovery mechanisms to ensure reliable pipeline execution.
- Troubleshoot pipeline failures, performance bottlenecks, data-quality issues, and production incidents.
- Optimize Spark workloads through effective use of partitioning, joins, caching, broadcast strategies, shuffle management, file sizing, and query execution plans.
- Develop reusable data engineering utilities, libraries, and frameworks using PySpark and Python.
- Write complex SQL queries for data transformation, reconciliation, validation, analysis, and reporting.
- Integrate Databricks pipelines with external applications and systems using REST APIs.
- Develop and support inbound and outbound file integrations using SFTP.
- Process structured and semi-structured data in formats such as CSV, JSON, Parquet, and Delta.
- Work with AWS services used in the data engineering ecosystem,



particularly Amazon S3.
- Implement exception handling, logging, auditing, monitoring, alerting, and operational recovery mechanisms.
- Implement data-quality validations and reconciliation controls across different pipeline stages.
- Participate in code reviews and ensure adherence to coding, security, performance, and data engineering best practices.
- Collaborate with architects, business analysts, source-system teams, QA teams, DevOps teams, and other engineering stakeholders.
- Support production deployments, release validation, operational monitoring, and troubleshooting of production data pipelines.
- Create and maintain technical documentation, pipeline specifications, operational procedures, and support runbooks.
- Use Cursor or similar agentic AI-enabled IDEs to improve development productivity, including:

- Generating and refactoring PySpark, Python, and SQL code
- Creating unit tests and data-validation scripts
- Troubleshooting errors and performance issues
- Generating technical documentation and code explanations
- Reviewing and validating AI-generated code for correctness, security, maintainability, and performance

Must-Have Skills

- Strong hands-on experience in Databricks-based data engineering.
- Advanced proficiency in PySpark, Spark SQL, and SQL.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Hands-on experience building ETL/ELT pipelines at enterprise scale.
- Strong experience with Delta Lake and Medallion Architecture.
- Practical experience with Delta Lake features such as MERGE, schema evolution, Time Travel, OPTIMIZE, and VACUUM.
- Hands-on experience developing and managing Databricks Workflows and Jobs.
- Strong understanding of Spark performance optimization, including:
- Data partitioning
- Join strategies
- Caching and persistence




- Shuffle optimization
- Data skew handling
- Query execution plans
- Small-file management

- Solid SQL skills, including complex joins, window functions, aggregations, reconciliation, and data-validation queries.
- Working knowledge of Python for developing reusable utilities, frameworks, and automation scripts.
- Experience working with CSV, JSON, Parquet, and Delta file formats.
- Hands-on experience integrating data pipelines with REST APIs and SFTP systems.
- Experience working with Amazon S3 in a data engineering environment.
- Experience implementing logging, exception handling, auditing, monitoring, alerting, and restart/recovery mechanisms.
- Experience troubleshooting and supporting production data pipelines.
- Hands-on knowledge of Cursor, GitHub Copilot, Windsurf, or another agentic AI-enabled development environment.
- Ability to write effective prompts, review AI-generated code, identify hallucinated or incorrect logic, and apply organizational security and data-privacy standards while using AI tools.
- Strong problem-solving, debugging, communication, and stakeholder-collaboration skills.

Preferred Skills

- Knowledge of Databricks Unity Catalog, access controls, lineage, and data governance.
- Experience with Databricks Asset Bundles, Databricks CLI, or CI/CD-based deployment approaches.
- Basic understanding of AWS IAM, Secrets Manager, CloudWatch, and Lambda.
- Experience with cloud security, secrets management, and credential-handling best practices.
- Knowledge of automated testing frameworks for PySpark and data pipelines.
- Familiarity with Git-based development, branching strategies, pull requests, and CI/CD pipelines.
- Experience working in Agile/Scrum delivery environments.
- Databricks or AWS certification would be an advantage.

Education and Experience

- Bachelor?s degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Relevant professional experience in data engineering, including substantial hands-on experience with Databricks, PySpark, SQL, and Delta Lake.
- Experience delivering and supporting enterprise-scale data platforms in production environments.

📌 DATA ENGINEER - Databricks (Bengaluru)
🏢 Happiest Minds Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - databricks (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - databricks (bengaluru) / bengaluru