19 Sep
|
Trigyn Technologies
|
India
19 Sep
Trigyn Technologies
India
Analytics Engineer-Senior AI / RAG Platform Engineer Position Id: G0926-0034
Job Type: Full Time
Country: India
Location: Delhi
Pay Rate: Open
Contact Recruiter: 912261400909
Experience Level: 5 Years
Work Arrangement: Hybrid / Remote Flexible
Executive Overview We are looking for an experienced, hands-on Senior AI & RAG Platform Engineer with 5 years of software engineering and applied AI experience to lead the architecture, configuration, optimisation, and production deployment of our generative AI capabilities. In this role, you will be the technical owner responsible for designing advanced Retrieval-Augmented Generation (RAG) pipelines, managing vector indexing and retrieval lifecycle, orchestrating foundation models (both open-source and proprietary), and embedding these AI services natively into our in-house custom-developed platform.
Key Responsibilities
1. Advanced RAG Architecture & Pipeline Configuration
- End-to-End Retrieval Pipelines: Design, configure, and optimise enterprise RAG pipelines combining dense vector search, sparse keyword search (BM25), reciprocal rank fusion (RRF), and cross-encoder reranking models (Cohere Rerank, BGE).
- Document Parsing & ETL: Build robust ingestion workflows for parsing, segmenting, and embedding multi-format documents (PDFs, Markdown, relational data, APIs) with semantic chunking and hierarchy preservation.
- Vector Lifecycle & Operations: Configure and maintain production vector databases (e.g., pgvector, Qdrant, Pinecone, Milvus, Weaviate), including multi-tenant partitioning, payload indexing, and HNSW index tuning.
- AI Model Selection, Orchestration & Agents
- Model Integration: Evaluate, select, and configure foundation models (OpenAI GPT-4o, Anthropic Claude 3.5, Llama 3.1/3.3, Mistral) based on cost, context limits, latency, and task complexity.
- Agentic Workflows & Tool Calling: Implement structured schema decoding (JSON/Pydantic), function calling,
and multi-step agentic execution flows (using LangGraph, LlamaIndex, or custom workflows).
- Fine-Tuning & Serving (Optional/Targeted): Conduct parameter-efficient fine-tuning (PEFT / LoRA / QLoRA) when necessary, and deploy open-weight models via inference servers such as vLLM, TensorRT-LLM, or Ollama.
- Custom Platform Integration & Engineering
- API & Microservice Development: Design and expose scalable REST, gRPC, and WebSocket streaming endpoints using Python (FastAPI, asyncio) to interface between LLM workflows and our proprietary platform microservices.
- State & Memory Management: Build sliding-window conversational memory, persistent session tracking, semantic caching (Redis / GPTCache), and tenant-level access control at the retrieval boundary.
- Async & Event-Driven Processing: Integrate heavy inference and document embedding pipelines into distributed message brokers and queues (Celery, Kafka, RabbitMQ, SQS).
- LLMOps, Guardrails & Quality Assurance
- Quantitative Evaluation: Establish continuous RAG evaluation frameworks (Ragas, TruLens, DeepEval) tracking Faithfulness, Answer Relevance, Context Precision, and Hallucination rates.
- Guardrails & Security: Implement strict guardrails (NeMo Guardrails, Llama Guard), prompt-injection mitigation, PII masking, and data privacy governance.
- Observability: Track latency, token utilization, rate limits, and end-to-end trace flows using Langfuse, Arize Phoenix, LangSmith, or OpenTelemetry.
Candidate Requirements & Qualifications Must-Have Requirements
- 5 years of professional backend/software engineering experience, with primary expertise in Python (FastAPI, Pydantic, asyncio, multiprocessing).
- 2 years of hands-on experience designing and operating RAG pipelines and production LLM-based solutions.
- Deep practical knowledge of AI orchestration frameworks: LlamaIndex, LangChain, LangGraph, or custom agent engines.
- Production experience with vector databases (e.g., pgvector, Qdrant, Pinecone, Weaviate, Milvus).
- Solid background in microservices architecture, RESTful/gRPC design, WebSocket streaming for token generation, and distributed caching (Redis).
- Experience with relational databases (PostgreSQL) and query performance tuning.
- Proficient in Docker, container orchestration (Kubernetes), CI/CD, and major cloud ecosystems (AWS, GCP, or Azure).
- Familiarity with RAG evaluation metrics and debugging hallucination/grounding issues.
Preferred / Nice-to-Have
- Experience deploying high-concurrency LLM inference backends using vLLM or TensorRT-LLM.
- Experience with Row-Level Security (RLS) and Role-Based Access Control (RBAC) applied to vector knowledge bases.
- Background in embedding model domain adaptation or LoRA fine-tuning.
- Prior experience working within custom SaaS or internal enterprise platforms.
Target Technical Stack Domain Technologies
- Languages & Frameworks: Python 3.11 , FastAPI, Pydantic, asyncio, Node.js (secondary)
- RAG & Orchestration: LlamaIndex, LangChain, LangGraph, DSPy, Unstructured.io
- Vector DBs & Search: pgvector (PostgreSQL), Qdrant, Pinecone, Weaviate, BM25 / OpenSearch
- LLMs & Embeddings: OpenAI, Anthropic Claude, Llama 3.1/3.3, Mistral, BGE / Voyage Embeddings
- LLMOps & Evaluation: Ragas, TruLens, Langfuse, Arize Phoenix, NeMo Guardrails
- Infrastructure: Docker, Kubernetes, Redis, RabbitMQ/Kafka, AWS/GCP, Terraform
📌 Analytics Engineer-Senior AI / RAG Platform Engineer (India)
🏢 Trigyn Technologies
📍 India