17 Sep
|
Intellistaff
|
Bengaluru
17 Sep
Intellistaff
Bengaluru
Data Engineer Azure Databricks
Experience: 5-12 Years
Location: Bangalore / Chennai / Hyderabad / Pune
Work Model: Hybrid
Role Type: Data Engineering / Azure / Databricks
Role Overview
We are looking for a hands-on Data Engineer with strong experience in Python, PySpark, Apache Spark, Azure Databricks and Advanced SQL.
The ideal candidate should have strong data engineering depth and engineering thinking, with the ability to independently develop, debug, optimize and support enterprise-scale data pipelines.
This is a hands-on engineering role. Candidates with primarily theoretical, support-only or certification-based exposure to Databricks/Azure will not be considered.
Key Responsibilities
- Design, develop and maintain scalable ETL/ELT data pipelines using Azure Databricks, PySpark, Python and SQL.
- Build data ingestion, transformation, validation, reconciliation and processing frameworks.
- Work with structured and semi-structured data across Oracle, SQL Server/Azure SQL and cloud data platforms.
- Develop and optimize PySpark/Spark jobs for performance and scalability.
- Work with Delta Lake, Lakehouse and Medallion Architecture.
- Implement incremental data processing, CDC, SCD, MERGE/UPSERT and watermark-based loading patterns.
- Work with Azure Data Factory, ADLS, Azure SQL and Synapse.
- Troubleshoot pipeline failures, data-quality issues and production incidents.
- Perform root-cause analysis and implement permanent fixes.
- Optimize Spark workloads involving partitions, joins, shuffles, data skew and query performance.
- Follow Git, CI/CD, coding standards and data governance practices.
- Collaborate with architects, business analysts and engineering teams to convert business requirements into technical solutions.
Mandatory Skills
- Python Solid
- Advanced SQL
- PySpark – Strong
- Apache Spark – Strong
- Azure Databricks – Hands-on
- Azure Data Factory
- ADLS
- Azure SQL / Synapse
- ETL/ELT
- Data pipeline development and debugging
- Spark performance tuning
- Data modelling fundamentals
Technical Expectations SkillMinimum ExpectationRecruiter Should VerifyPythonStrongCan independently write and debug codeSQLAdvancedCTEs, window functions, joins and query logicPySparkStrongTransformations, joins and aggregationsSparkStrongExecution, optimization, partitions and shufflesDatabricksHands-onJobs, clusters, notebooks and DeltaAzureHands-onADF, ADLS, Azure SQL/SynapseETLStrongEnd-to-end pipeline ownershipDebuggingCriticalCan troubleshoot failed code/pipelinesPerformance TuningCriticalCan explain an actual optimization performedData ModellingGoodFacts, dimensions, SCD and incremental loadsBFSI/InsurancePreferredInsurance data/process exposure
Good to Have
- Delta Lake / Unity Catalog
- Azure Synapse
- Spark performance tuning
- Data Lakehouse / Medallion Architecture
- Kafka / Event Hub
- Azure DevOps / CI-CD
- Oracle
- Data archival / retention / purge
- Insurance / BFSI experience
Candidate Profile
Candidates should demonstrate
- Strong hands-on Python development.
- Real project experience with Databricks and PySpark.
- Ability to explain the complete data pipeline architecture they have worked on.
- Ability to write and explain SQL rather than only use pre-written queries.
- Strong understanding of Spark execution and optimization.
- Experience troubleshooting production issues independently.
- Experience owning data pipelines from ingestion through downstream consumption.
Candidates who only mention Azure/Databricks in their resume without project-level implementation experience should not be considered.
📌 Azure Data Engineer (Bengaluru)
🏢 Intellistaff
📍 Bengaluru