We're hiring 3 Data Engineers to build and scale an enterprise Cloud Data Lakehouse on Azure Databricks + ADLS Gen2 . You'll own pipelines end-to-end — from raw ingestion through Bronze/Silver/Gold — and work directly with data scientists, analytics engineers, and business stakeholders.
This is a hands-on engineering role, not a notebook-only role. If you write modular, tested, version-controlled Python and care about what a job costs to run, you'll fit.
What you'll do
- Build scalable batch and streaming pipelines with PySpark, Spark SQL, and Databricks
- Orchestrate workflows via Azure Data Factory, Databricks Workflows, and Delta Live Tables
- Implement Medallion Architecture on Delta Lake — partitioning, schema enforcement, data quality expectations
- Tune clusters, jobs, and queries to kill bottlenecks and cut cloud spend
- Enforce row/column-level access and lineage through Unity Catalog
- Ship via Azure DevOps or GitHub Actions with real CI/CD and test coverage
- Deliver clean Gold-layer datasets for ML models and Power BI
What we're looking for (3–5 years)
- Hands-on Azure Databricks, Apache Spark (PySpark/Spark SQL), Delta Lake
- Azure core: ADLS Gen2, ADF, Key Vault
- Strong Python + advanced SQL (window functions, complex joins, query tuning)
- Dimensional modelling (Star/Snowflake) and Medallion Lakehouse patterns
- Git, code reviews, modular Python outside notebooks, basic CI/CD
- Bachelor's in CS/IT or equivalent practical experience
- Explicit communication with non-technical stakeholders
Nice to have Unity Catalog · Kafka or Azure Event Hubs · Terraform/Bicep · Databricks Certified Data Engineer Associate
Details
Remote (India) · Full-time · 3 positions - Immediate joiner only To apply: [Apply here or
[email protected] ] — subject line "Azure DE — [Your Name]"
📌 Azure Data Engineer (Databricks) — Remote, India
🏢 Zoft AI
📍 India