17 Sep
|
TechBlocks
|
Hyderabad
17 Sep
TechBlocks
Hyderabad
Role Overview
We are seeking an experienced Data Engineer with strong hands-on expertise in Databricks, Delta Lake, Databricks SQL, Workflows, Unity Catalog, advanced SQL, and Python. The candidate will be responsible for designing, developing, and operationalizing production-grade data pipelines, with a strong focus on integrating data from REST APIs and building scalable, reliable data solutions.
The ideal candidate should have experience developing enterprise data pipelines, implementing dimensional data models, establishing data quality frameworks, and working with governed data environments. Experience with the Azure data ecosystem and AI/engineering-tool telemetry will be an added advantage.
Key Responsibilities
Databricks Data Engineering
- - Design, develop, and maintain production-grade data pipelines using Databricks.
- Develop scalable data processing solutions using Delta Lake and Databricks SQL.
- Build and manage Databricks Workflows for pipeline orchestration, scheduling, monitoring, and dependency management.
- Implement reliable and reusable data ingestion and transformation frameworks.
- Optimize Databricks workloads for performance, scalability, reliability, and cost efficiency.
- Implement appropriate error handling, logging, monitoring, and recovery mechanisms for production pipelines.
REST API Data Integration
- - Develop data pipelines that ingest data from REST APIs and external enterprise systems.
- Design reusable API ingestion frameworks capable of handling authentication, pagination, rate limits, retries, incremental extraction, and error handling.
- Transform API responses into structured datasets suitable for downstream analytics.
- Implement mechanisms for incremental and historical data ingestion.
- Troubleshoot API connectivity, data availability, schema changes, and ingestion failures.
Data Modelling & SQL
- - Design and implement dimensional data models for analytical workloads.
- Develop fact and dimension tables and establish appropriate relationships for reporting and analytics.
- Write advanced SQL for data transformation, validation, aggregation, and analytical processing.
- Optimize complex SQL queries and Databricks workloads for performance.
- Ensure data models are scalable, maintainable, and aligned with business reporting requirements.
Python Development
- - Develop robust data engineering applications and pipeline components using Python.
- Build reusable Python libraries and utilities for ingestion, transformation, validation, and automation.
- Implement exception handling, logging, configuration management, and testing practices.
- Use Python to automate operational and data engineering activities.
Data Quality & Reliability
- - Design and implement data quality frameworks across ingestion and transformation pipelines.
- Establish automated checks for data completeness, accuracy, consistency, uniqueness, and validity.
- Implement data reconciliation and validation mechanisms between source systems and target datasets.
- Monitor pipeline failures and data-quality issues and drive timely resolution.
- Establish quality thresholds, alerts, and exception-handling mechanisms for production data pipelines.
Unity Catalog & Data Governance
- - Configure and manage data assets within Databricks Unity Catalog.
- Implement appropriate access controls and permissions for data, schemas, tables, and other governed assets.
- Support data governance, security, discoverability, and controlled access across the Databricks environment.
- Maintain appropriate metadata and documentation for data assets.
- Support governance and lineage requirements across data pipelines.
Production Operations
- - Deploy and support Databricks pipelines in production environments.
- Monitor pipeline execution, performance, failures, and data quality.
- Troubleshoot production incidents and perform root-cause analysis.
- Implement CI/CD and deployment practices for data engineering workloads where applicable. Collaborate with Data Architects, Data Analysts, DevOps, Product Owners,
and business stakeholders.
Required Skills & Experience
- - Strong hands-on experience as a Data Engineer working with Databricks in production environments.
- Strong expertise in:
- Delta Lake
- Databricks SQL
- Databricks Workflows
- Unity Catalog
- Strong experience developing production-grade data pipelines using REST APIs.
- Advanced proficiency in SQL, including complex joins, CTEs, window functions, aggregations, and query optimization.
- Strong programming experience with Python for data engineering and pipeline development.
- Hands-on experience with dimensional data modelling, including fact and dimension design.
- Experience designing and implementing data quality frameworks and automated data validation.
- Robust understanding of data ingestion, transformation, orchestration, and production operations.
- Experience troubleshooting data pipeline failures, data discrepancies, performance issues, and integration problems.
- Good understanding of data security, access controls, and governance within enterprise data platforms. Ability to develop scalable and maintainable data engineering solutions in an enterprise environment.
Strongly Preferred Skills
- - Experience implementing Medallion Architecture using Bronze, Silver, and Gold data layers.
- Hands-on experience with Delta Live Tables (DLT).
- Strong experience with Unity Catalog governance, data lineage, and metadata management.
- Experience designing adapter or canonical schema patterns for integrating data from multiple external systems.
- Experience with the Azure data stack, including services such as Azure Data Factory, Azure Data Lake Storage, Azure Functions, and Azure Synapse.
- Experience integrating engineering and software-development data sources into Databricks.
- Experience working with AI/engineering-tool usage and cost telemetry.
- Experience ingesting and analyzing telemetry from AI coding assistants and engineering productivity platforms.
- Knowledge of cloud-based CI/CD and DevOps practices for Databricks deployments.
- Experience with Git and Azure DevOps.
- Exposure to data observability and advanced data-quality tooling.
📌 Data Engineer (Hyderabad)
🏢 TechBlocks
📍 Hyderabad