29 Aug
|
Trigent Software - Professional Services
|
India
29 Aug
Trigent Software - Professional Services
India
Role Overview
We are seeking a Senior Data Engineer specializing in Python, Unstructured NLP/Document Ingestion, and Neo4j Graph Data Modeling. You will lead the design and implementation of end-to-end ETL pipelines that extract data from enterprise sources (SharePoint, REST APIs, SQL, flat files), perform text manipulation and vectorization, and structure complex output into JSON payload routines optimized for downstream Neo4j ingestion.
Key Responsibilities
- Unstructured Ingestion &
- Extraction:
Build scalable pipelines using Python and PySpark to ingest and parse complex unstructured documents (PDFs, Word docs, SharePoint sites, APIs).
- NLP &
- Text Processing:
Apply NLP techniques to clean, extract, and manipulate natural language text, preparing entity-relationship metadata and generating vector embeddings.
- JSON Schema Strategy:
Structure, validate, and optimize rich, nested JSON payloads tailored specifically for graph data representation and downstream ETL routines.
- Graph Modeling &
- Neo4j Integration:
Collaborate with Neo4j engineers to design graph models, optimize Cypher queries, and build seamless graph ingestion routines.
- Vector Store Integration:
Interface with vector databases or vector store extensions to support hybrid semantic search and GraphRAG architectures.
Required Qualifications
- Core Tech Stack:
Solid proficiency in Python, PySpark, SQL, and JSON manipulation.
- Graph Database Expertise:
Hands-on experience with Neo4j, Cypher querying, and graph data modeling.
- Document Parsing &
- Connectors:
Proven track record ingesting and processing unstructured documents via enterprise connectors (SharePoint APIs, REST APIs, blob storage).
- NLP &
- Text Processing:
Hands-on experience using NLP libraries (e.g., SpaCy, NLTK, Hugging Face, LangChain) for entity extraction, text chunking, and metadata generation.
- Vector Stores:
Exposure to vector databases (e.g., Pinecone, Milvus, Qdrant, Chroma) or Neo4j vector indexes.
Nice to Have
- Experience building Knowledge Graphs or GraphRAG (Retrieval-Augmented Generation) applications.
- Production experience with enterprise cloud data platforms (AWS, Azure, or GCP).
📌 Senior Data Engineer – Neo4j (India)
🏢 Trigent Software - Professional Services
📍 India