AI Data Engineer (Uttar Pradesh)

AI Data Engineer (Uttar Pradesh)

30 Jul
|
Sparix Global
|
Uttar Pradesh

30 Jul

Sparix Global

Uttar Pradesh

Position: AI Data Engineer

Level: Senior

Experience: 8-10 years

Location: Noida (hybrid)

Budget: around 15-20 LPA

About the role

You own the data layer that makes our AI products actually work. Every RAG system we ship lives or dies on what you build the ingestion pipelines that pull from messy enterprise sources, the embedding and chunking choices that decide whether retrieval is precise or noisy, the vector infrastructure that has to stay fast and fresh in production, and the OCR/NLP pre-processing that handles real-world Indian enterprise data (Hinglish, regional scripts, scanned PDFs, semi-structured forms).

This is not a traditional data warehousing role. You are building AI-native data infrastructure vector databases, embedding pipelines, retrieval evaluation harnesses, document ingestion at scale for clients in BFSI, Healthcare, Manufacturing, and Retail/D2C.

Required qualifications

- 8-10 years in data engineering, with the last 2+ years building data infrastructure for LLM/RAG applications in production (not just analytics or BI pipelines).
- Solid Python (pandas, PySpark, FastAPI for data services) and strong SQL.
- Production experience with at least two vector databases Pinecone, Weaviate, pgvector, Qdrant, Milvus, or Chroma.
- Hands-on experience with embedding models OpenAI embeddings, Cohere, BGE family, sentence-transformers, or equivalent including knowing when each is the right choice.




- Production experience with at least one orchestration framework Airflow, Dagster, Prefect, or equivalent.
- Hands-on OCR/document AI experience AWS Textract, Azure Document Intelligence, Google Document AI, Tesseract, PaddleOCR, or Unstructured.io.
- Cloud experience (AWS, Azure, or GCP) covering storage, compute, and managed data services.
- Demonstrable understanding of retrieval evaluation you can explain why a RAG system is failing using metrics, not anecdotes.

Preferred qualifications
- Experience with multilingual NLP Indic languages (Hindi, Tamil, Telugu, Bengali, etc.), Hinglish, transliteration handling, IndicBERT/IndicTrans, AI4Bharat models.

- Experience with hybrid search architectures BM25 + dense retrieval, reranking with Cohere Rerank or cross-encoders.

- PII detection and masking tooling Presidio, AWS Comprehend PII, regex+NER hybrid systems.

- Experience with knowledge graphs (Neo4j, ArangoDB) and graph-augmented RAG.

- Familiarity with DPDP Act 2023 and HIPAA data handling requirements.

- Prior consulting or client-facing delivery experience.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 AI Data Engineer (Uttar Pradesh)
🏢 Sparix Global
📍 Uttar Pradesh

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai data engineer (uttar pradesh) / uttar pradesh

Subscribe to this job alert:

Get the latest job offers by email for: ai data engineer (uttar pradesh) / uttar pradesh