Key Responsibilities
Databricks Platform Engineering
Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.
Delta Lake & Lakehouse Architecture
Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
Manage schema evolution, table versioning,
time travel, and Change Data Feed (CDF) for incremental processing.
Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL)
Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and energetic partition pruning.
Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.
MLflow & AI/ML Workloads
Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
Suppor
📌 Data Engineer (India)
🏢 EXL
📍 India