06 Sep
|
TalentOla
|
Chennai
We are seeking a highly skilled Data Lead with 8–10 years of experience in Big Data and Spark technologies. The ideal candidate will have hands-on experience with Databricks and a robust product-building mindset. You will lead the validation of scalable data pipelines, ensuring data integrity across Delta Lake and Unity Catalog environments
Key Responsibilities
Utilize Databricks Notebooks to author comprehensive technical test cases and validation logic using PySpark and Spark SQL.
Design and execute testing strategies for both structured and unstructured data (PDFs, Text files, and RTF documents) to ensure high-fidelity transformation into structured formats within the Data Lakehouse.
Validate all Source-to-Target (S2T) mappings across bronze, silver, and gold layers to ensure data lineage and integrity.
Leverage strong PYSpark/ Spark SQL knowledge to create mockup data for exhaustive edge-case testing.
Design, implement, and maintain an automated regression suite to ensure pipeline stability across code releases
Validate end-to-end data accuracy from the Lakehouse layers to final analytical dashboards and reporting tools.
Lead the review of test cases and drive development best practices, including code reviews and performance optimization for test scripts.
Implement and verify data quality checks within the Unity Catalog to ensure proper data lineage and access control
Partner with data engineers and analysts to align quality benchmarks with Healthix business goals.
Required Skills & Experience:
Experience in Big Data Engineering using Apache Spark and related technologies.
Experience with Databricks (Notebooks, Jobs, Workflows, Delta Lake).
Minimum 2–3 years of hands-on experience with Databricks, including:
Delta Lake for scalable and reliable data lake architectures.
Unity Catalog for centralized data governance.
Job & Workflow orchestration, including DLT pipelines.
Experience validating DLT and modern ETL/ELT design patterns.
Hands-on experience with workflow orchestration using Databricks Workflows
Proficient in PySpark, SparkSQL, Advanced SQL and Spark optimization techniques.
Experience with AWS cloud platform.
Excellent communication, stakeholder management, and leadership skills.
📌 Lead Databricks Engineer (Chennai)
🏢 TalentOla
📍 Chennai