Role Purpose Build the Entity Graph. This role converts the resolved, canonical entity data into a production graph — implementing node and edge projections, loading and optimising the graph, enabling multi-hop traversal, and validating that the graph faithfully represents entity, ownership and affiliation relationships back to source. Key Responsibilities Graph schema implementation — implement the graph schema defined by the architect: node types, edge types, properties, keys and constraints. Relational-to-graph projection — build and operate the projection logic that converts canonical relational entity tables into graph nodes and edges, with full source attribution. Graph loading & pipelines — develop and optimise graph load processes; implement incremental re-projection triggered by change data capture. Relationship modelling — implement entity-to-entity, ownership (parent/subsidiary), officer/director and affiliation relationships, including bitemporal handling where required. Traversal & query engineering — write and optimise graph queries (GQL/Cypher-style) supporting multi-hop traversal scenarios such as corporate family trees and shared-agent affiliations. Performance optimisation — tune graph queries and load performance; profile traversal cost; address scale bottlenecks. Graph quality validation — implement automated checks on node/edge counts, orphan detection, relationship integrity and source traceability. Consumption support — support exposure of graph data through API and SQL endpoints, and collaborate on GraphRAG indexing over graph projections. Required Skills & Experience Skill Area Specific Requirements Graph Databases Hands-on with one or more of: Fabric Graph, Neo4j, Cosmos DB (Gremlin), TigerGraph, Neptune. Solid LPG modelling Graph Query GQL (ISO/IEC 39075),
Cypher or Gremlin; multi-hop traversal, path queries, pattern matching, query optimisation Data Engineering Python/PySpark, SQL, Delta Lake, ETL/ELT pipeline development, incremental/CDC processing Microsoft Fabric Lakehouse, OneLake, Spark notebooks, Data Factory pipelines, SQL analytics endpoint Modelling Converting relational schemas to graph models, key/edge design, handling many-to-many and hierarchical structures Quality & Ops Graph validation techniques, monitoring, troubleshooting load failures, documentation Must-Have Qualifications 7 years data engineering with 3 years hands-on graph database development Proven experience modelling and loading a production graph from relational sources Strong graph query language proficiency (GQL, Cypher or Gremlin) Solid PySpark and SQL engineering skills Experience with hierarchical/ownership data structures and recursive relationships Nice-to-Have Microsoft Fabric Graph experience (native LPG on OneLake) Exposure to GraphRAG or graph-based retrieval Experience in corporate entity, KYC, fraud-network or supply-chain graph domains Graph algorithms (community detection, centrality, shortest path) Key Deliverables Owned Implemented graph schema and node/edge projection logic Baseline entity graph populated in Fabric Incremental re-projection on CDC Validated multi-hop traversal scenarios Graph quality validation checks and performance tuning results Dual Role / Complementary Skills Highly complementary with the VectorDB Engineer (Role 4) — GraphRAG requires graph traversal and vector retrieval working together. If consolidating headcount, these two roles can be merged into a single "Graph & Vector Engineer", since the vector workload is concentrated in Phase 2 while graph schema work runs earlier. Also cross-covers with the Data Engineer on Spark-based pipeline work.
📌 Graph DB Engineer (Haryana)
🏢 EXL
📍 Haryana