31 Aug
|
Finarb
|
Kolkata
Job Description
Data QA Engineer
/n
Location: Kolkata, India (Onsite/Hybrid)
/n
Experience: 2–4 years
/n
Employment Type: Full-time
/n
ABOUT THE ROLE
/n
We are hiring a Data QA Engineer for our data engineering team. This role is responsible for
/n
validating ETL pipelines and data platforms built on Azure Data Factory, Databricks, and
/n
Microsoft Fabric, along with the downstream tables, reports, and models they feed. The core
/n
responsibility is verifying data correctness, which is distinct from confirming that a pipeline
/n
executed without errors — the two are frequently conflated, and this role exists to keep them
/n
separate.
/n
KEY RESPONSIBILITIES
/n
/n
- Design and execute test plans covering source-to-target validation and transformation logic
/n
- for ETL pipelines
/n
- Write SQL and PySpark scripts to verify accuracy, completeness, and consistency in Delta
/n
- Lake tables
/n
- Perform regression testing on every pipeline change; a successful pipeline run does not
/n
- guarantee correct output, and validation must be independent of execution status
/n
- Build and maintain reusable data quality checks (e.g., Great Expectations, dbt tests, or
/n
- custom PySpark frameworks) in place of one-off manual queries
/n
- Reconcile data between source systems and target Lakehouse/Warehouse layers, and
/n
- investigate root cause when discrepancies are found
/n
- Validate schema conformance, null/duplicate handling, referential integrity, and business
/n
- rule adherence across Bronze/Silver/Gold layers
/n
- Validate Microsoft Fabric artifacts — Lakehouses, Warehouses, and semantic models —
/n
- including Direct Lake mode behavior and cross-domain data consistency
/n
- Apply consistent QA methodology across tools; the underlying platform (Databricks, Fabric,
/n
- or otherwise) should not change how rigorously data is validated
/n
- Document test cases and defects with enough detail for engineers to act on them without
/n
- requiring additional clarification
/n
- Work directly with the data engineering team on requirement clarification and defect
/n
- resolution
/n
/n
REQUIRED SKILLS
/n
/n
- Robust SQL, with the ability to write validation queries independently
/n
- Working proficiency in Python/PySpark for scripting data checks
/n
- Solid understanding of ETL concepts: staging, transformations, incremental loads, SCD
/n
- handling
/n
- Hands-on experience with Azure Data Factory and Databricks
/n
- Working knowledge of Microsoft Fabric (Lakehouse, Warehouse, semantic models)
/n
- Familiarity with Delta Lake and medallion architecture
/n
- Ability to read transformation logic and determine expected output
/n
- General data QA methodology that transfers across tools and platforms, not skills tied to a
/n
- single vendor stack
/n
- Experience with defect tracking and structured test documentation (Jira, Test Rail, or
/n
- equivalent) PREFERRED QUALIFICATIONS
/n
- Experience with a data quality framework (Great Expectations, Deequ, dbt tests)
/n
- Familiarity with Unity Catalog and general data governance/lineage tooling
/n
- Experience validating pipelines in a regulated domain (pharma, healthcare, finance) where
/n
- lineage and auditability are requirements
/n
- Exposure to CI/CD for test automation
/n
- Azure, Databricks, or Fabric certification
/n
/n
QUALIFICATIONS
/n
Bachelor’s degree in Computer Science, IT, or a related field
/n
2–4 years of experience in data QA, data validation, or ETL testing; manual/UI testing
/n
experience without pipeline exposure does not meet this requirement
/n
CANDIDATE FIT
/n
This role requires the ability to determine why a data discrepancy occurred, not simply flag
/n
that one exists. Candidates whose QA background is primarily manual/UI testing and who are
/n
seeking to transition into data-focused work should not apply for this position; the required
/n
data depth is expected from day one.
/n
📌 Data Quality Assurance Lead (Kolkata)
🏢 Finarb
📍 Kolkata