Role Purpose
Build the Entity Graph. This role converts the resolved, canonical entity data into a production graph — implementing node and edge projections, loading and optimising the graph, enabling multi-hop traversal, and validating that the graph faithfully represents entity, ownership and affiliation relationships back to source.
Key Responsibilities
- Graph schema implementation — implement the graph schema defined by the architect: node types, edge types, properties, keys and constraints.
- Relational-to-graph projection — build and operate the projection logic that converts canonical relational entity tables into graph nodes and edges, with full source attribution.
- Graph loading & pipelines — develop and optimise graph load processes; implement incremental re-projection triggered by change data capture.
- Relationship modelling — implement entity-to-entity, ownership (parent/subsidiary), officer/director and affiliation relationships, including bitemporal handling where required.
- Traversal & query engineering — write and optimise graph queries (GQL/Cypher-style) supporting multi-hop traversal scenarios such as corporate family trees and shared-agent affiliations.
- Performance optimisation — tune graph queries and load performance; profile traversal cost; address scale bottlenecks.
- Graph quality validation — implement automated checks on node/edge counts, orphan detection, relationship integrity and source traceability.
- Consumption support — support exposure of graph data through API and SQL endpoints, and collaborate on GraphRAG indexing over graph projections.
Required Skills & Experience
Skill Area
Specific Requirements
Graph Databases
Hands-on with one or more of: Fabric Graph, Neo4j, Cosmos DB (Gremlin), TigerGraph, Neptune. Strong LPG modelling
Graph Query
GQL (ISO/IEC 39075),
Cypher or Gremlin; multi-hop traversal, path queries, pattern matching, query optimisation
Data Engineering
Python/PySpark, SQL, Delta Lake, ETL/ELT pipeline development, incremental/CDC processing
Microsoft Fabric
Lakehouse, OneLake, Spark notebooks, Data Factory pipelines, SQL analytics endpoint
Modelling
Converting relational schemas to graph models, key/edge design, handling many-to-many and hierarchical structures
Quality & Ops
Graph validation techniques, monitoring, troubleshooting load failures, documentation
Must-Have Qualifications
- 7+ years data engineering with 3+ years hands-on graph database development
- Proven experience modelling and loading a production graph from relational sources
- Robust graph query language proficiency (GQL, Cypher or Gremlin)
- Solid PySpark and SQL engineering skills
- Experience with hierarchical/ownership data structures and recursive relationships
Nice-to-Have
- Microsoft Fabric Graph experience (native LPG on OneLake)
- Exposure to GraphRAG or graph-based retrieval
- Experience in corporate entity, KYC, fraud-network or supply-chain graph domains
- Graph algorithms (community detection, centrality, shortest path)
Key Deliverables Owned
- Implemented graph schema and node/edge projection logic
- Baseline entity graph populated in Fabric
- Incremental re-projection on CDC
- Validated multi-hop traversal scenarios
- Graph quality validation checks and performance tuning results
Dual Role / Complementary Skills
Highly complementary with the VectorDB Engineer (Role 4) — GraphRAG requires graph traversal and vector retrieval working together. If consolidating headcount, these two roles can be merged into a single Graph & Vector Engineer, since the vector workload is concentrated in Phase 2 while graph schema work runs earlier. Also cross-covers with the Data Engineer on Spark-based pipeline work.
📌 Graph DB Engineer (Gurugram)
🏢 EXL
📍 Gurugram