19 Aug
|
Godrej Properties
|
Mumbai
19 Aug
Godrej Properties
Mumbai
Role & responsibilities
- Build and maintain connectors from enterprise systems (Salesforce, SharePoint, S3, SQL DBs) to the AI retrieval layer. Handle incremental sync and change detection.
- Design and optimise document chunking strategies for different content types (PDFs, HTML, structured tables, multi-modal). Maintain preprocessing pipelines.
- Select, fine-tune, and maintain embedding models (Titan Embeddings, penAI Ada, BGE, multilingual models). Benchmark quality on domain data.
- Manage OpenSearch Serverless / Pinecone / pgvector clusters: indexing, refresh cycles, cost optimisation, namespace management.
- Implement hybrid search (dense + sparse), re-ranking, query expansion, and retrieval evaluation. Own the precision/recall benchmarks.
- Enforce document-level access controls in retrieval, ensure PII redaction pipelines,
and maintain data lineage for all AI-consumed datasets.
Preferred candidate profile
- 4+ years in data engineering (Python, PySpark, or SQL-heavy roles)
- Hands-on experience building RAG pipelines: chunking, embedding, vector retrieval
- Proficiency with at least one vector database: OpenSearch, Pinecone, Weaviate, or pgvector
- Solid SQL and ETL pipeline experience; familiarity with Airflow or Glue
- Understanding of embedding models when to use domain-specific vs general-purpose
- Experience with enterprise document formats: PDF extraction, SharePoint, Confluence, structured
📌 Artificial Intelligence Engineer (Mumbai)
🏢 Godrej Properties
📍 Mumbai