13 Sep
|
Finarb
|
Kolkata
Data QA Engineer /n Location : Kolkata, India (Onsite/Hybrid) /n Experience : 2–4 years /n Employment Type: Full-time /n ABOUT THE ROLE /n We are hiring a Data QA Engineer for our data engineering team. This role is responsible for /n validating ETL pipelines and data platforms built on Azure Data Factory, Databricks, and /n Microsoft Fabric, along with the downstream tables, reports, and models they feed. The core /n responsibility is verifying data correctness, which is distinct from confirming that a pipeline /n executed without errors — the two are frequently conflated, and this role exists to keep them /n separate. /n KEY RESPONSIBILITIES /n /n
- Design and execute test plans covering source-to-target validation and transformation logic
/n
- for ETL pipelines
/n
- Write SQL and PySpark scripts to verify accuracy, completeness, and consistency in Delta
/n
- Lake tables
/n
- Perform regression testing on every pipeline change; a successful pipeline run does not
/n
- guarantee correct output, and validation must be independent of execution status
/n
- Build and maintain reusable data quality checks (e.g., Outstanding Expectations, dbt tests, or
/n
- custom PySpark frameworks) in place of one-off manual queries
/n
- Reconcile data between source systems and target Lakehouse/Warehouse layers, and
/n
- investigate root cause when discrepancies are found
/n
- Validate schema conformance, null/duplicate handling, referential integrity, and business
/n
- rule adherence across Bronze/Silver/Gold layers
/n
- Validate Microsoft Fabric artifacts — Lakehouses, Warehouses, and semantic models —
/n
- including DirectLake mode behavior and cross-domain data consistency
/n
- Apply consistent QA methodology across tools; the underlying platform (Databricks, Fabric,
/n
- or otherwise) should not change how rigorously data is validated
/n
- Document test cases and defects with enough detail for engineers to act on them without
/n
- requiring additional clarification
/n
- Work directly with the data engineering team on requirement clarification and defect
/n
- resolution
/n /n REQUIRED SKILLS /n /n
- Strong SQL, with the ability to write validation queries independently
/n
- Working proficiency in Python/PySpark for scripting data checks
/n
- Solid understanding of ETL concepts: staging, transformations, incremental loads, SCD
/n
- handling
/n
- Hands-on experience with Azure Data Factory and Databricks
/n
- Working knowledge of Microsoft Fabric (Lakehouse, Warehouse, semantic models)
/n
- Familiarity with Delta Lake and medallion architecture
/n
- Ability to read transformation logic and determine expected output
/n
- General data QA methodology that transfers across tools and platforms, not skills tied to a
/n
- single vendor stack
/n
- Experience with defect tracking and structured test documentation (Jira, TestRail, or
/n
- equivalent) PREFERRED QUALIFICATIONS
/n
- Experience with a data quality framework (Great Expectations, Deequ, dbt tests)
/n
- Familiarity with Unity Catalog and general data governance/lineage tooling
/n
- Experience validating pipelines in a regulated domain (pharma, healthcare, finance) where
/n
- lineage and auditability are requirements
/n
- Exposure to CI/CD for test automation
/n
- Azure, Databricks, or Fabric certification
/n /n QUALIFICATIONS /n Bachelor’s degree in Computer Science, IT, or a related field /n 2–4 years of experience in data QA, data validation, or ETL testing; manual/UI testing /n experience without pipeline exposure does not meet this requirement /n CANDIDATE FIT /n This role requires the ability to determine why a data discrepancy occurred, not simply flag /n that one exists. Candidates whose QA background is primarily manual/UI testing and who are /n seeking to transition into data-focused work should not apply for this position; the required /n data depth is expected from day one. /n
📌 Data Quality Assurance Lead (Kolkata)
🏢 Finarb
📍 Kolkata