Data Engineering
- Design, develop, and maintain scalable data pipelines using Databricks.
- Build end-to-end ETL/ELT solutions using PySpark, Python, and SQL.
- Develop reliable batch and streaming data pipelines.
- Optimize Spark jobs and Databricks workloads for performance and cost.
- Implement Delta Lake best practices and Medallion Architecture.
- Develop reusable frameworks and maintain high engineering standards.
- Refactor legacy ETL solutions into up-to-date Databricks-based architectures.
- Ensure production stability through monitoring, alerting, and troubleshooting.
Technical Leadership
- Define technical standards and best practices.
- Review solution design and code quality.
- Collaborate with Architects, Product Owners, Business Teams, and Platform Teams.
- Mentor junior engineers and provide technical guidance.
- Drive continuous improvement in engineering practices.
Delivery & Stakeholder Management
- Own delivery of data engineering initiatives.
- Manage project timelines, sprint planning, and technical risks.
- Communicate project status with stakeholders.
- Ensure production readiness and successful releases.
Leadership (Preferred)
- Lead and mentor Data Engineering teams.
- Support hiring, onboarding, and capability development.
- Conduct technical reviews and mentor engineers.
- Experience managing engineering teams is highly desirable but not mandatory.
Preferred candidate profile
Databricks
- Databricks Workspace
- Delta Lake
- Databricks Workflows / Jobs
- Auto Loader
- Structured Streaming
- Unity Catalog
- Delta Live Tables (DLT)
Programming
- PySpark
- Python
- SQL
- Spark SQL
Data Engineering
- ETL / ELT
- Data Pipeline Development
- Data Warehousing
- Data Modelling
- Batch & Streaming Pipelines
- Performance Optimization
Cloud
- Azure or AWS (GCP is a plus)
- ADLS / S3 DevOps
- Git
- CI/CD
- Azure DevOps / GitHub