18 Sep
|
Qentelli
|
India
Role summary
Build and operationalize the Fabric pipeline designed by the architect: ingestion from multiple sources, Spark-based transformation logic, Gold layer tables (both MLV and DirectQuery-served), and ongoing maintenance/optimization through go-live and stabilization. This is a sustained build-and-run role, following an existing architecture rather than designing it from scratch.
Key responsibilities
Build and maintain Spark notebooks for BronzeSilver transformations, including parsing semi-structured (JSON) and NoSQL-sourced data
Implement entity-resolution/join logic across sources (structured relational + semi-structured + NoSQL)
Build and schedule Materialized Lake Views and/or DirectQuery views per the architecture's real-time/analytical split
Set up and maintain Data Pipeline orchestration (ingestion, transform, refresh scheduling, dependency chains)
Monitor and tune performance: table optimization/vacuuming, small-file compaction, Direct Lake troubleshooting
Maintain and extend the mirroring/ingestion connections (Oracle mirroring, OFSC API polling, MongoDB connector) as data volume or requirements evolve
Support integration testing with the AI agent layer (Fabric Data Agent and/or direct SQL/MCP tool queries)
Document pipeline logic and hand over runbooks for BAU support
Must-have skills
Robust hands-on PySpark and Spark SQL comfortable writing production transformation logic, not just tutorials
Working experience with Microsoft Fabric: Lakehouse, notebooks, Data Pipelines, Delta tables
Experience with Delta Lake concepts: schema evolution, optimize/vacuum, change data feed
SQL proficiency (T-SQL for Warehouse/SQL analytics endpoint work)
Experience with API-based data ingestion (REST APIs, JSON parsing, pagination/polling patterns)
Comfortable working from an architecture document and translating design decisions into working code
Nice-to-have
Materialized Lake Views experience specifically
Experience with MongoDB or other NoSQL Spark connectors
Power BI / semantic model exposure (helpful for testing Gold layer outputs, not core to the role)
DP-203 or DP-700 (Fabric Data Engineer) certification.
📌 Microsoft Fabric Data Engineer Hyderabad (India)
🏢 Qentelli
📍 India