15 Sep
|
Dexian
|
Hyderabad
Job Description
Role - Data Engineer – Data Pipelines & Governance
n
Location: Chennai (Preferred) / Pune / Hyderabad / Bangalore (Onsite/Remote)
n
Contract
n
Experience: 6+ years in data engineering, with robust Azure and Databricks experience
n
Job Description: Data Engineer – Data Pipelines & Governance
n
We are seeking a hands-on Data Engineer to develop, optimize, and maintain automated data pipelines supporting data governance and analytics initiatives. This role will focus on building production-ready workflows for ingestion, transformation, quality checks, lineage capture, access auditing, cost usage analysis, retention tracking, and metadata integration, primarily using Azure Databricks, Azure Data Lake, and MS Purview.
n
Key Responsibilities
n
n
- Pipeline Development – Design, build, and deploy robust ETL/ELT pipelines in Databricks (PySpark, SQL, Delta Lake) to ingest, transform, and curate governance and operational metadata from multiple sources landed in Databricks.
n
- Granular Data Quality Capture – Implement profiling logic to capture issue-level metadata (source table, column, timestamp, severity, rule type) to support drill-down from dashboards into specific records and enable targeted remediation.
n
- Governance Metrics Automation – Develop data pipelines to generate metrics for dashboards covering data quality, lineage, job monitoring, access & permissions, query cost, usage & consumption, retention & lifecycle, policy enforcement, sensitive data mapping, and governance KPIs.
n
- MS Purview Integration – Automate asset onboarding, metadata enrichment, classification tagging,
and lineage extraction for integration into governance reporting.
n
- Data Retention & Policy Enforcement – Implement logic for retention tracking and policy compliance monitoring (masking, RLS, exceptions).
n
- Job & Query Monitoring – Build pipelines to track job performance, SLA adherence, and query costs for cost and performance optimization.
n
- Metadata Storage & Optimization – Maintain curated Delta tables for governance metrics, structured for efficient dashboard consumption.
n
- Testing & Troubleshooting – Monitor pipeline execution, optimize performance, and resolve issues quickly.
n
- Collaboration – Work closely with the lead engineer, QA, and reporting teams to validate metrics and resolve data quality issues.
n
- Security & Compliance – Ensure all pipelines meet organizational governance, privacy, and security standards.
n
n
Required Qualifications
n
n
- Bachelor's degree in computer science, Engineering, Information Systems, or related field
n
- 6+ years of hands-on data engineering experience, with Azure Databricks and Azure Data Lake
n
- Proficiency in PySpark, SQL, and ETL/ELT pipeline design
n
- Demonstrated experience building granular data quality checks and integrating governance logic into pipelines
n
- Working knowledge of MS Purview for metadata management, lineage capture, and classification
n
- Experience with Azure Data Factory or equivalent orchestration tools
n
- Understanding of data modeling, metadata structures, and data cataloging concepts
n
- Strong debugging, performance tuning, and problem-solving skills
n
- Ability to document pipeline logic and collaborate with cross-functional teams
n
📌 Azure Data Engineer (Hyderabad)
🏢 Dexian
📍 Hyderabad