06 Aug
|
Yinghuali Automotive Interiors
|
Pune
06 Aug
Yinghuali Automotive Interiors
Pune
Job Position: AI Engineer
Experience: Minimum 4 to 5 years of relevant software engineering experience, with a meaningful and demonstrable portion specifically on retrieval, search or RAG systems
Probation Period: 6 Months
Location:Pune, India
Function: Engineering / AI Platform
Employment type: Permanent, full-time
Role purpose:
CortexMI is CT Automotives AI-powered engineering and quality intelligence platform, built on Anthropics Claude API with a retrieval-augmented generation (RAG) architecture. The platform ingests complex automotive quality documents PFMEAs, Control Plans, 8Ds and ECRs and makes them queryable in natural language by engineers across the UK, China, Mexico and Turkey.
The integrity of the platform depends entirely on the retrieval pipeline: documents must be extracted cleanly, chunked intelligently, embedded, stored and retrieved with complete fidelity. This role owns that pipeline end to end from document upload through to the chunks that reach the language model at query time.
This is a deep specialist role, not a generalist backend position. We are looking for someone who has built and debugged production RAG systems and understands their failure modes intimately. The successful candidate will be the person the team turns to when a document does not retrieve correctly, when uploads are slow, or when a query returns incomplete results and who can diagnose the fault to a specific layer and fix it with evidence.
The pipeline you will own
The CortexMI ingestion and retrieval pipeline runs as follows, and you will be accountable for every stage of it:
Azure Blob Storage document extraction chunking embedding Qdrant vector storage retrieval at query time Claude API.
A concrete sense of the data: a single AIAG-VDA 2019 PFMEA has 24 columns, uses merged cells for station and process groupings, and an individual populated failure-mode row can run to 100150 tokens. A dense process station can contain well over a hundred such rows. Extraction, chunking and retrieval must handle this structure without losing a single row because in a quality system, a dropped failure mode is a real and unacceptable gap.
Core responsibilities:
- Own the full ingestion and retrieval pipeline end to end: upload to Azure Blob Storage,
text and table extraction, chunking, embedding, vector storage in Qdrant, and retrieval at query time.
- Solve structured-document extraction problems specific to engineering data in particular the correct resolution of merged cells in Excel-based PFMEAs and Control Plans, where station and process values span multiple rows and must be propagated down correctly before chunking. Incorrect extraction here silently corrupts everything downstream.
- Design and tune the chunking strategy for dense, tabular, multi-column documents: setting chunk size and overlap so that large stations are not fragmented into excessive chunks, and so that later chunks in a large station are never dropped during embedding or upsert.
- Optimise embedding and vector upsert through batching collecting all chunks for a document, embedding them in batched API calls, and upserting to Qdrant in batches rather than sequentially to eliminate dropped chunks and reduce upload latency.
- Tune retrieval quality: top-K configuration, payload/metadata filtering, and document-scoped retrieval, so that a query reliably returns all relevant chunks rather than truncating before the later sections of a document.
- Build verification and observability into the pipeline: row-count validation against the source document, chunk-completeness checks, extraction confidence reporting, and clear stage-by-stage logging so any failure can be diagnosed to a specific layer.
- Drive upload latency down from the current multi-minute range to seconds, through batching and asynchronous processing.
- Work directly with the Lead AI Engineer and CT quality engineers to validate retrieval accuracy against known-good source documents before any live data goes into the system.
Essential skills and experience:
- Demonstrable hands-on experience building production RAG / retrieval pipelines not prototypes or coursework.
Candidates should be able to describe specific retrieval failures they have personally diagnosed and fixed.
- Deep working knowledge of vector databases, ideally Qdrant collections, payload and metadata filtering, upsert behaviour, and retrieval tuning.
- Robust Python, including document-processing libraries openpyxl or equivalent for Excel structure and merged-cell handling, and robust PDF and table extraction tooling.
- A practical understanding of chunking strategies for structured and tabular data: embedding models, token budgeting, and the trade-offs between chunk size, overlap, retrieval precision and recall.
- Experience with embedding APIs, including batching, rate limits, and timeout and retry handling.
- Cloud storage and pipeline experience Azure Blob Storage or equivalent (AWS S3, GCS) and comfort building asynchronous, batched processing pipelines.
- A rigorous, diagnostic mindset: able to isolate whether a fault sits in extraction, chunking, embedding, storage or retrieval, and to prove a fix with evidence rather than assumption.
Desirable:
- Experience with Anthropics Claude API or comparable LLM APIs used in a RAG context.
- Exposure to manufacturing, automotive or engineering quality documentation (PFMEA, FMEA, Control Plans, APQP) or a demonstrated ability to learn a complex technical domain quickly.
- Familiarity with voice and multilingual pipelines (speech-to-text, text-to-speech, automatic language detection).
What success looks like in the first 90 days
- Documents of any supported type upload and fully index in under 30 seconds.
- Every row and section of an uploaded document is verifiably present in the vector store, confirmed by an automated row-count check against the source file.
- Merged-cell resolution is reliable across all PFMEA and Control Plan formats, with no 0 rows metadata failures.
- A retrieval query for any station or section returns the complete, correct content demonstrated across the full six-question validation set on a known-good document.
- Pipeline observability is in place: each stage logs clearly and a failure can be traced to a single layer within minutes.
4th June 2026
CT Automotive India
📌 RAG Engineer (Pune)
🏢 Yinghuali Automotive Interiors
📍 Pune