24 Sep
|
Tiger Analytics
|
Bengaluru
24 Sep
Tiger Analytics
Bengaluru
Role & responsibilities
We are looking for an experienced Azure Databricks Data Engineer to build and optimize scalable cloud-based data engineering solutions on Microsoft Azure. The ideal candidate should have strong hands-on experience in Azure Databricks, PySpark, Azure Data Factory (ADF), Spark SQL, Delta Lake, and ADLS Gen2 with expertise in designing enterprise-grade ETL/ELT pipelines and modern Lakehouse architectures.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Azure Databricks, PySpark, Spark SQL, and Azure Data Factory (ADF).
- Build and implement Lakehouse/Medallion Architecture (Bronze, Silver, Gold layers).
- Develop ETL/ELT pipelines for batch data processing and transformation.
- Perform data cleansing, deduplication, incremental loads, and implement SCD Type 1 & Type 2 using Delta Lake.
- Optimize Spark applications by tuning partitions, caching, memory management, and Adaptive Query Execution (AQE).
- Handle performance optimization using Broadcast Joins, Data Skew handling, Partitioning, and File Optimization techniques.
- Monitor and troubleshoot production jobs using Spark UI, Azure Monitor, and Databricks job logs.
- Build CI/CD pipelines using Git, Azure DevOps, and Databricks Asset Bundles (DABs).
- Collaborate with Data Architects, Analysts, and Business teams to deliver high-quality data solutions.
- Follow best practices in coding standards, testing, deployment, and documentation.
Required Skills
- 3+ years (Data Engineer) / 5+ years (Senior Data Engineer) in Azure Data Engineering.
- Strong experience with
- Azure Databricks
- PySpark
- Spark SQL
- Azure Data Factory (ADF)
- Azure Data Lake Storage Gen2 (ADLS)
- Delta Lake
- Hands-on experience with:
- Delta MERGE
- Time Travel
- VACUUM
- Managed & External Tables
- Robust SQL skills including
- Window Functions
- ROW_NUMBER
- RANK
- DENSE_RANK
- LEAD
- LAG
- Experience in
- ETL/ELT Development
- Data Modeling
- Performance Tuning
- Data Quality
- Schema Evolution
- Good understanding of Git, Azure DevOps, CI/CD, and production deployments.
Preferred Skills
- Structured Streaming
- Unity Catalog
- Infrastructure as Code (Terraform/Bicep)
- Azure Monitor
- Data Governance
- Python
- pytest / Unit Testing
- Cost Optimization in Databricks
- Data Quality Frameworks
Mandatory Skills
- Azure Databricks
- PySpark
- Azure Data Factory (ADF)
- Spark SQL
- ADLS Gen2
- Delta Lake
- Azure Data Engineering
- SQL
- ETL
- Azure Cloud
Good to Have
- Unity Catalog
- Structured Streaming
- Azure DevOps
- Git
- Databricks Asset Bundles
- Lakehouse Architecture
- Medallion Architecture
- Data Modeling
Preferred candidate profile
📌 Azure Databricks Data Engineer / Senior Data Engineer (Bengaluru)
🏢 Tiger Analytics
📍 Bengaluru