24 Sep
|
Happiest Minds Technologies
|
Bengaluru
24 Sep
Happiest Minds Technologies
Bengaluru
Senior Databricks Data Engineer
Experience Required
68 years of overall experience in data engineering, including strong hands-on experience with Databricks, PySpark, Spark SQL, SQL, and Delta Lake.
Role Overview
We are looking for a Senior Databricks Data Engineer to design, develop, test, and maintain scalable ETL/ELT data pipelines. The candidate will be responsible for hands-on development, technical-design support, code reviews, performance optimization, production troubleshooting, and mentoring junior developers.
The role requires close collaboration with technical leads, architects, business analysts, QA teams, source-system teams, and other engineering stakeholders. The candidate should be able to independently develop complex data pipelines and contribute to the continuous improvement of engineering standards and reusable frameworks.
The candidate should also have practical knowledge of Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled development tools.
Key Responsibilities
Data Pipeline Development
- Design, develop, test, and maintain scalable ETL/ELT data pipelines using Databricks and PySpark.
- Develop and optimize complex data-transformation logic using PySpark, Spark SQL, and SQL.
- Build and maintain Delta Lake tables following Bronze, Silver, and Gold layers of the Medallion Architecture.
- Implement Delta Lake capabilities, including:
- MERGE and upsert operations
- Schema enforcement and schema evolution
- Time Travel and data-version management
- OPTIMIZE and VACUUM
- Incremental and change-based data processing
- Develop reusable data engineering components, utilities, and frameworks using PySpark and Python.
- Write complex SQL queries for data transformation, reconciliation, validation, troubleshooting, and analysis.
- Process structured and semi-structured data in CSV, JSON, Parquet, and Delta formats.
Technical-Design Support
- Work with technical leads and architects to understand solution architecture and translate technical designs into development components.
- Contribute to low-level design documents, data-flow diagrams, source-to-target mappings, and interface specifications.
- Participate in technical discussions related to data models, pipeline patterns, integrations, performance, scalability, and maintainability.
- Identify technical dependencies, risks, assumptions, and data-quality concerns during design and development.
- Provide development estimates and support sprint and release planning.
- Recommend appropriate technical solutions for assigned modules and pipeline components.
Databricks Workflow and Integration Development
- Develop and manage Databricks Workflows and Jobs for pipeline orchestration and scheduling.
- Configure task dependencies, parameters, retries, notifications, and failure-handling mechanisms.
- Integrate Databricks pipelines with external applications and source systems using REST APIs.
- Develop and support inbound and outbound file integrations using SFTP.
- Work with Amazon S3 for reading, writing, storing, and managing data used by Databricks pipelines.
- Apply secure methods for handling API credentials, SFTP keys, secrets, and other sensitive configuration information.
Performance Optimization
- Analyse and optimize Spark workloads using appropriate:
- Data-partitioning strategies
- Join and broadcast strategies
- Caching and persistence
- Shuffle optimization
- Data-skew handling
- Query execution plans
- File-compaction and small-file management techniques
- Troubleshoot performance bottlenecks in PySpark, Spark SQL, SQL, and Delta Lake pipelines.
- Ensure that pipelines are scalable, efficient, reliable, and maintainable.
Code Review and Engineering Quality
- Participate in peer code reviews for PySpark, Python, Spark SQL, and SQL components.
- Review code for functional correctness, performance, maintainability, reusability, security, and adherence to coding standards.
- Ensure that developed components include proper exception handling, logging, auditing, monitoring, and recovery mechanisms.
- Develop and maintain unit tests, data-quality checks, reconciliation controls, and validation scripts.
- Follow established Git branching, pull-request, documentation, and deployment practices.
- Identify opportunities for code refactoring, framework improvements, and technical-debt reduction.
- Ensure that code-review comments are addressed before deployment.
Mentoring and Team Collaboration
- Provide technical guidance and development support to junior data engineers.
- Mentor junior developers in Databricks, PySpark, Spark SQL, Delta Lake, Python, SQL, and data-engineering best practices.
- Conduct knowledge-sharing sessions, technical walkthroughs, and pair-programming activities when required.
- Help junior developers troubleshoot development, performance, testing, and production issues.
- Collaborate with architects, technical leads, business analysts, QA teams, DevOps teams, source-system teams, and other engineering stakeholders.
- Communicate technical issues, risks, dependencies, and progress clearly to the Technical Lead or Project Manager.
Production Support
- Support production deployments, release validation, and post-deployment monitoring.
- Troubleshoot pipeline failures, data-quality issues, performance problems, and production incidents.
- Perform root-cause analysis and implement permanent corrective and preventive solutions.
- Develop restart, retry, recovery, and reconciliation mechanisms for failed or partially completed pipeline executions.
- Maintain technical documentation, troubleshooting guides, operational procedures, and support runbooks.
Agentic AIEnabled Development
- Use Cursor, GitHub Copilot, Windsurf, or similar agentic AI-enabled IDEs to improve development productivity.
- Use AI development tools for:
- Generating and refactoring PySpark, Python, and SQL code
- Creating unit tests and data-validation scripts
- Debugging pipeline failures
- Identifying potential performance improvements
- Generating technical documentation and code explanations
- Write effective prompts and provide relevant technical context to AI development tools.
- Review AI-generated code for correctness, security, performance, maintainability, and compliance with coding standards.
- Ensure that confidential data, credentials, and proprietary information are handled according to organizational security and AI-usage policies.
Must-Have Skills
- 68 years of overall experience in data engineering.
- Solid hands-on experience with Databricks, PySpark, Spark SQL, and SQL.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Experience designing and developing enterprise-scale ETL/ELT data pipelines.
- Strong working knowledge of Delta Lake and Medallion Architecture.
- Hands-on experience with:
- Delta MERGE operations
- Schema enforcement and evolution
- Time Travel
- OPTIMIZE and VACUUM
- Incremental data processing
- Experience developing and managing Databricks Workflows and Jobs.
- Strong knowledge of Spark performance optimization, including partitioning, joins, caching, shuffle operations, data skew, and query execution plans.
- Strong SQL skills, including complex joins, window functions, aggregations, reconciliation, and data-validation queries.
- Working knowledge of Python for developing reusable utilities and automation scripts.
- Experience integrating data pipelines with REST APIs and SFTP.
- Experience working with CSV, JSON, Parquet, and Delta file formats.
- Hands-on experience working with Amazon S3.
- Experience implementing logging, exception handling, auditing, monitoring, alerting, and restart/recovery mechanisms.
- Experience supporting production data pipelines and resolving technical issues.
- Experience participating in code reviews and following engineering best practices.
- Ability to independently develop complex data pipelines with limited supervision.
- Ability to provide technical guidance and mentoring to junior developers.
- Practical knowledge of Cursor, GitHub Copilot, Windsurf, or another agentic AI-enabled development tool.
- Strong analytical, problem-solving, debugging, communication, and collaboration skills.
Preferred Skills
- Experience with Databricks Unity Catalog, data lineage, access controls, and governance.
- Experience with Databricks Asset Bundles, Databricks CLI, or CI/CD-based deployment.
- Basic knowledge of AWS IAM, Secrets Manager, CloudWatch, and Lambda.
- Experience with Git-based development, branching strategies, pull requests, and CI/CD pipelines.
- Knowledge of automated testing frameworks for PySpark and data pipelines.
- Understanding of cloud security and secrets-management best practices.
- Experience working in an Agile/Scrum delivery environment.
- Databricks or AWS certification would be an advantage.
Education Bachelors degree in Computer Science, Information Technology, Engineering, or a related discipline.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr DATA ENGINEER - Databricks (Bengaluru)
🏢 Happiest Minds Technologies
📍 Bengaluru