Key Responsibilities...
- Migrate SQL Server Stored Procedures to Databricks Notebooks, leveraging PySpark and Spark SQL for complex transformations.
- Design, build, and maintain incremental data load pipelines to handle dynamic updates from various sources, ensuring scalability and efficiency.
- Develop robust data ingestion pipelines to load data into the Databricks Bronze layer from relational databases, APIs, and file systems.
- Implement incremental data transformation workflows to update silver and gold layer datasets in near real-time, adhering to Delta Lake best practices.
- Integrate Airflow with Databricks to orchestrate end-to-end workflows, including dependency management, error handling, and scheduling.
- Understand business and technical requirements, translating them into scalable Databricks solutions.
- Optimize Spark jobs and queries for performance, scalability, and cost-efficiency in a distributed environment.
- Implement robust data quality checks,
monitoring solutions, and governance frameworks within Databricks.
- Collaborate with team members on Databricks best practices, reusable solutions, and incremental loading strategies.
All you need is...
- Bachelor's degree in computer science, Information Systems, or a related discipline.
- 4+ years of hands-on experience with Databricks, including expertise in Databricks SQL, PySpark, and Spark SQL. (Must)
- Proven experience in incremental data loading techniques into Databricks, leveraging Delta Lake's features (e.g., time travel, MERGE INTO).
- Solid understanding of data warehousing concepts, including data partitioning, and indexing for efficient querying.
- Proficiency in T-SQL and experience in migrating SQL Server Stored Procedures to Databricks.
📌 Consultant | Data Engineer (Pune)
🏢 amdocs
📍 Pune