31 Jul
|
Finarb
|
Kolkata
Data QA Engineer
Location: Kolkata, India (Onsite/Hybrid)
Experience: 2–4 years
Employment Type: Full-time
ABOUT THE ROLE
We are hiring a Data QA Engineer for our data engineering team. This role is responsible for validating ETL pipelines and data platforms built on Azure Data Factory, Databricks, and
Microsoft Fabric, along with the downstream tables, reports, and models they feed. The core responsibility is verifying data correctness, which is distinct from confirming that a pipeline executed without errors — the two are frequently conflated, and this role exists to keep them separate.
KEY RESPONSIBILITIES
Design and execute test plans covering source-to-target validation and transformation logic for ETL pipelines
Write SQL and PySpark scripts to verify accuracy, completeness, and consistency in Delta
Lake tables
Perform regression testing on every pipeline change; a successful pipeline run does not guarantee correct output, and validation must be independent of execution status
Build and maintain reusable data quality checks (e.g., Great Expectations, dbt tests, or custom PySpark frameworks) in place of one-off manual queries
Reconcile data between source systems and target Lakehouse/Warehouse layers, and investigate root cause when discrepancies are found
Validate schema conformance, null/duplicate handling, referential integrity, and business rule adherence across Bronze/Silver/Gold layers
Validate Microsoft Fabric artifacts — Lakehouses, Warehouses, and semantic models —
including DirectLake mode behavior and cross-domain data consistency
Apply consistent QA methodology across tools; the underlying platform (Databricks, Fabric,
or otherwise) should not change how rigorously data is validated
Document test cases and defects with enough detail for engineers to act on them without requiring additional clarification
Work directly with the data engineering team on requirement clarification and defect resolution
REQUIRED SKILLS
Robust SQL, with the ability to write validation queries independently
Working proficiency in Python/PySpark for scripting data checks
Solid understanding of ETL concepts: staging, transformations, incremental loads, SCD handling
Hands-on experience with Azure Data Factory and Databricks
Working knowledge of Microsoft Fabric (Lakehouse, Warehouse, semantic models)
Familiarity with Delta Lake and medallion architecture
Ability to read transformation logic and determine expected output
General data QA methodology that transfers across tools and platforms, not skills tied to a single vendor stack
Experience with defect tracking and structured test documentation (Jira, TestRail, or equivalent)PREFERRED QUALIFICATIONS
Experience with a data quality framework (Great Expectations, Deequ, dbt tests)
Familiarity with Unity Catalog and general data governance/lineage tooling
Experience validating pipelines in a regulated domain (pharma, healthcare, finance) where lineage and auditability are requirements
Exposure to CI/CD for test automation
Azure, Databricks, or Fabric certification
QUALIFICATIONS
Bachelor’s degree in Computer Science, IT, or a related field
2–4 years of experience in data QA, data validation, or ETL testing; manual/UI testing experience without pipeline exposure does not meet this requirement
CANDIDATE FIT
This role requires the ability to determine why a data discrepancy occurred, not simply flag that one exists. Candidates whose QA background is primarily manual/UI testing and who are seeking to transition into data-focused work should not apply for this position; the required data depth is expected from day one.
📌 Data Quality Assurance Lead (Kolkata)
🏢 Finarb
📍 Kolkata