11 Sep
|
Quantiphi
|
Bengaluru
11 Sep
Quantiphi
Bengaluru
Job Description
Summary
n
We are looking for an experienced Senior Data Engineer to join our team at Quantiphi, a leader in AI-first engineering and innovative GenAI solutions. In this pivotal role, you will design and architect robust data pipelines that power advanced AI applications, specifically focusing on building high-performance ingestion engines for unstructured data, vector databases, and scalable real-time streaming architectures. You will bridge the gap between complex data engineering and modern AI, ensuring data quality, lineage, and observability as we deploy cutting-edge Agentic AI solutions.
n
Responsibilities
n
n
- Design and build batch ETL pipelines that ingest large unstructured document corpora - handling PDF, RTF, and JSON file formats, consuming pre-extracted text and metadata, applying document chunking strategies, and loading processed outputs into a vector store and analytical platforms.
n
- Design partitioning, parallelism, and throughput strategies to sustain high-volume ingestion.
n
- Build real-time streaming ingestion pipelines consuming document creation and update events, processing them through chunking and embedding generation.
n
- Build vector and indexing pipelines targeting vector store platforms (e.g., Vertex AI Vector Search, Pinecone, Weaviate, Milvus, pgvector)
n
- Manage vector store index operations: bulk load, incremental upsert, index refresh, and consistency validation post-ingestion.
n
- Implement metadata tagging on vector records to support filtered retrieval at query time.
n
- Implement end-to-end lineage tracking for vector records - linking each chunk back to its source document.
n
- Implement document-level change tracking to detect upstream updates, deletions,
and re-ingestion events and trigger appropriate vector store operations (upsert, delete, re-index) without full corpus re-processing.
n
- Tag all vector records with pipeline run metadata - ingestion timestamp, pipeline version, chunking strategy version, embedding model version, to support retrieval quality debugging and model refresh traceability.
n
n
Must Have Skills
n
n
- Data Engineering - Experience building production-grade batch and streaming data pipelines supporting enterprise analytics and AI workloads.
n
- Python- Strong proficiency in Python for data processing, pipeline development, automation, and transformation frameworks.
n
- Streaming Technologies - Experience building real-time ingestion pipelines using Apache Kafka, GCP Pub/Sub, Azure Event Hubs, or equivalent event streaming platforms.
n
- Data Platforms - Experience working with BigQuery/Synapse, Databricks, Snowflake, or equivalent cloud-native analytical platforms.
n
- Data Modeling & SQL - Robust SQL skills and experience designing scalable analytical and operational data models.
n
- Vector databases - Familiarity with Pinecone, Vertex AI Vector Search, Weaviate, Milvus, pgvector, or similar retrieval platforms.
n
- Document processing - Familiarity with PDF, RTF, JSON, OCR pipelines, and metadata extraction workflows.
n
- Data Quality & Governance - Experience implementing validation frameworks, reconciliation, lineage, metadata management, monitoring, and pipeline observability.
n
- Cloud Platforms - Hands-on experience building cloud-native data solutions using GCS/ADLS, BigQuery/Synapse, Dataflow/ADF, Cloud Composer/Airflow, or equivalent services.
n
- Regulated Data - Experience working with regulated data environments where security, auditability, governance, and compliance are critical.
n
📌 Senior Data Engineer(15-30 days joiners only) (Bengaluru)
🏢 Quantiphi
📍 Bengaluru