08 Sep
|
Happiest Minds Technologies
|
Bengaluru
08 Sep
Happiest Minds Technologies
Bengaluru
: Senior Databricks Data Engineer
Role Overview
We are looking for an experienced Senior Databricks Data Engineer to design, develop, and maintain scalable ETL/ELT data pipelines using Databricks, PySpark, Spark SQL, and Delta Lake. The candidate should have strong hands-on experience in data transformation, pipeline orchestration, performance optimization, data quality, and production support within an AWS-based data engineering environment.
The role also requires practical experience using Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools to accelerate development, testing, debugging, documentation, and code optimization.
Key Responsibilities
- Design, develop, test, and maintain scalable ETL/ELT data pipelines using Databricks and PySpark.
- Develop and optimize data transformation logic using PySpark, Spark SQL, and SQL.
- Build and maintain Delta Lake tables following Bronze, Silver, and Gold layers of the Medallion Architecture.
- Implement Delta Lake capabilities, including:
- MERGE and upsert operations
- Schema enforcement and schema evolution
- Time Travel and version management
- OPTIMIZE and VACUUM
- Incremental and change-based data processing
- Develop and manage Databricks Workflows and Jobs for pipeline orchestration, scheduling, dependency management, retries, and alerting.
- Design productive incremental-load and restart/recovery mechanisms to ensure reliable pipeline execution.
- Troubleshoot pipeline failures, performance bottlenecks, data-quality issues, and production incidents.
- Optimize Spark workloads through effective use of partitioning, joins, caching, broadcast strategies, shuffle management, file sizing, and query execution plans.
- Develop reusable data engineering utilities, libraries, and frameworks using PySpark and Python.
- Write complex SQL queries for data transformation, reconciliation, validation, analysis, and reporting.
- Integrate Databricks pipelines with external applications and systems using REST APIs.
- Develop and support inbound and outbound file integrations using SFTP.
- Process structured and semi-structured data in formats such as CSV, JSON, Parquet, and Delta.
- Work with AWS services used in the data engineering ecosystem,
particularly Amazon S3.
- Implement exception handling, logging, auditing, monitoring, alerting, and operational recovery mechanisms.
- Implement data-quality validations and reconciliation controls across different pipeline stages.
- Participate in code reviews and ensure adherence to coding, security, performance, and data engineering best practices.
- Collaborate with architects, business analysts, source-system teams, QA teams, DevOps teams, and other engineering stakeholders.
- Support production deployments, release validation, operational monitoring, and troubleshooting of production data pipelines.
- Create and maintain technical documentation, pipeline specifications, operational procedures, and support runbooks.
- Use Cursor or similar agentic AI-enabled IDEs to improve development productivity, including:
- Generating and refactoring PySpark, Python, and SQL code
- Creating unit tests and data-validation scripts
- Troubleshooting errors and performance issues
- Generating technical documentation and code explanations
- Reviewing and validating AI-generated code for correctness, security, maintainability, and performance
Must-Have Skills
- Strong hands-on experience in Databricks-based data engineering.
- Advanced proficiency in PySpark, Spark SQL, and SQL.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Hands-on experience building ETL/ELT pipelines at enterprise scale.
- Strong experience with Delta Lake and Medallion Architecture.
- Practical experience with Delta Lake features such as MERGE, schema evolution, Time Travel, OPTIMIZE, and VACUUM.
- Hands-on experience developing and managing Databricks Workflows and Jobs.
- Strong understanding of Spark performance optimization, including:
- Data partitioning
- Join strategies
- Caching and persistence
- Shuffle optimization
- Data skew handling
- Query execution plans
- Small-file management
- Solid SQL skills, including complex joins, window functions, aggregations, reconciliation, and data-validation queries.
- Working knowledge of Python for developing reusable utilities, frameworks, and automation scripts.
- Experience working with CSV, JSON, Parquet, and Delta file formats.
- Hands-on experience integrating data pipelines with REST APIs and SFTP systems.
- Experience working with Amazon S3 in a data engineering environment.
- Experience implementing logging, exception handling, auditing, monitoring, alerting, and restart/recovery mechanisms.
- Experience troubleshooting and supporting production data pipelines.
- Hands-on knowledge of Cursor, GitHub Copilot, Windsurf, or another agentic AI-enabled development environment.
- Ability to write effective prompts, review AI-generated code, identify hallucinated or incorrect logic, and apply organizational security and data-privacy standards while using AI tools.
- Strong problem-solving, debugging, communication, and stakeholder-collaboration skills.
Preferred Skills
- Knowledge of Databricks Unity Catalog, access controls, lineage, and data governance.
- Experience with Databricks Asset Bundles, Databricks CLI, or CI/CD-based deployment approaches.
- Basic understanding of AWS IAM, Secrets Manager, CloudWatch, and Lambda.
- Experience with cloud security, secrets management, and credential-handling best practices.
- Knowledge of automated testing frameworks for PySpark and data pipelines.
- Familiarity with Git-based development, branching strategies, pull requests, and CI/CD pipelines.
- Experience working in Agile/Scrum delivery environments.
- Databricks or AWS certification would be an advantage.
Education and Experience
- Bachelor?s degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Relevant professional experience in data engineering, including substantial hands-on experience with Databricks, PySpark, SQL, and Delta Lake.
- Experience delivering and supporting enterprise-scale data platforms in production environments.
📌 DATA ENGINEER - Databricks (Bengaluru)
🏢 Happiest Minds Technologies
📍 Bengaluru