Role Purpose
nBuild the Entity Graph. This role converts the resolved, canonical entity data into a production graph — implementing node and edge projections, loading and optimising the graph, enabling multi-hop traversal, and validating that the graph faithfully represents entity, ownership and affiliation relationships back to source.
nKey Responsibilities
n
n
- Graph schema implementation — implement the graph schema defined by the architect: node types, edge types, properties, keys and constraints.
n
- Relational-to-graph projection — build and operate the projection logic that converts canonical relational entity tables into graph nodes and edges, with full source attribution.
n
- Graph loading & pipelines — develop and optimise graph load processes; implement incremental re-projection triggered by change data capture.
n
- Relationship modelling — implement entity-to-entity, ownership (parent/subsidiary), officer/director and affiliation relationships, including bitemporal handling where required.
n
- Traversal & query engineering — write and optimise graph queries (GQL/Cypher-style) supporting multi-hop traversal scenarios such as corporate family trees and shared-agent affiliations.
n
- Performance optimisation — tune graph queries and load performance; profile traversal cost; address scale bottlenecks.
n
- Graph quality validation — implement automated checks on node/edge counts, orphan detection, relationship integrity and source traceability.
n
- Consumption support — support exposure of graph data through API and SQL endpoints, and collaborate on GraphRAG indexing over graph projections.
n
nRequired Skills & Experience
nSkill Area
nSpecific Requirements
nGraph Databases
nHands-on with one or more of: Fabric Graph, Neo4j, Cosmos DB (Gremlin), TigerGraph, Neptune. Strong LPG modelling
nGraph Query
nGQL (ISO/IEC 39075),
Cypher or Gremlin; multi-hop traversal, path queries, pattern matching, query optimisation
nData Engineering
nPython/PySpark, SQL, Delta Lake, ETL/ELT pipeline development, incremental/CDC processing
nMicrosoft Fabric
nLakehouse, OneLake, Spark notebooks, Data Factory pipelines, SQL analytics endpoint
nModelling
nConverting relational schemas to graph models, key/edge design, handling many-to-many and hierarchical structures
nQuality & Ops
nGraph validation techniques, monitoring, troubleshooting load failures, documentation
nMust-Have Qualifications
n
n
- 7+ years data engineering with 3+ years hands-on graph database development
n
- Proven experience modelling and loading a production graph from relational sources
n
- Solid graph query language proficiency (GQL, Cypher or Gremlin)
n
- Solid PySpark and SQL engineering skills
n
- Experience with hierarchical/ownership data structures and recursive relationships
n
nNice-to-Have
n
n
- Microsoft Fabric Graph experience (native LPG on OneLake)
n
- Exposure to GraphRAG or graph-based retrieval
n
- Experience in corporate entity, KYC, fraud-network or supply-chain graph domains
n
- Graph algorithms (community detection, centrality, shortest path)
n
nKey Deliverables Owned
n
n
- Implemented graph schema and node/edge projection logic
n
- Baseline entity graph populated in Fabric
n
- Incremental re-projection on CDC
n
- Validated multi-hop traversal scenarios
n
- Graph quality validation checks and performance tuning results
n
nDual Role / Complementary Skills
nHighly complementary with the VectorDB Engineer (Role 4) — GraphRAG requires graph traversal and vector retrieval working together. If consolidating headcount, these two roles can be merged into a single "Graph & Vector Engineer", since the vector workload is concentrated in Phase 2 while graph schema work runs earlier. Also cross-covers with the Data Engineer on Spark-based pipeline work.
📌 Graph DB Engineer (Gurugram)
🏢 EXL
📍 Gurugram