08 Sep
|
CHRYSELYS
|
Chennai
Graph RAG - Knowledge Engineer
Design, build and operate the Python/FastAPI services that extract entities and relationships from unstructured documents, resolve them to canonical identifiers, maintain the knowledge graph, and serve graph-augmented retrieval alongside vector search for multi-hop and relational questions.
Responsibilities
- Build entity and relation extraction services over unstructured documents molecules, brands, indications, therapeutic areas, endpoints, claims.
- Build entity resolution: alias handling, blocking and candidate generation, fuzzy and embedding matching, calibrated thresholds, human review routing.
- Design and maintain the graph schema and ontology; incremental ingest, node and edge deduplication and merging, provenance on every edge.
- Fuse graph and vector results into a single ranked, cited context for the retrieval service.
- Instrument, monitor and support the services in production.
Qualifications
- 5–9 years software engineering, with demonstrable knowledge-graph construction and applied NLP delivered to production.
- Has built a knowledge graph from unstructured text — not queried an existing one, and not a CRUD application on a graph database.
- Graph at production scale.
Millions of nodes and edges; incremental updates with secure node identity; supernode and traversal-explosion handling with bounded depth and timeouts.
- Entity resolution at corpus scale. Blocking and candidate generation that avoid O(n) comparison, with measured precision on a labelled sample.
- Graph database in production. Neo4j, Amazon Neptune or equivalent; Cypher / openCypher fluency.
- Ontology and taxonomy modelling. Schema evolution without breaking downstream consumers; judgement on node vs. edge vs. property.
- Extraction. NER and relation extraction — LLM-based, model-based (spaCy, scispaCy, transformers) or hybrid, with the judgement to choose.
- Graph vs. vector judgement. Knows where graph retrieval wins — multi-hop, relational, comparative and aggregate questions — and that hybrid is the production norm.
- Python and FastAPI in production. Python 3.11+, async, Pydantic, Docker, pytest, Git and CI; AWS as a consumer (S3, ECS/EKS, Bedrock, Neptune or self-hosted Neo4j).
Preferred Biomedical ontologies and registries: UMLS, MeSH, SNOMED, RxNorm, ICD-10, DrugBank, ChEMBL.
- Life sciences or pharma domain experience; RDF/SPARQL alongside property graphs.
- GraphRAG approaches: community detection for corpus-level summarisation, local vs. global search.
📌 Graph RAG - Knowledge Engineer (Chennai)
🏢 CHRYSELYS
📍 Chennai