24 Sep
|
Ramp infotech
|
Noida
24 Sep
Ramp infotech
Noida
AI Engineer (LLM / RAG / Document Intelligence)
Role Summary
We are seeking an AI Engineer to build the AI core of a new B2B platform being delivered to a client. Client and project details are confidential and will be disclosed at onboarding under a confidentiality agreement.
The systems operate in a domain where an incorrectly extracted value is a serious defect: accuracy, grounding and controllability take priority over raw capability. The role covers production LLM engineering end-to-end: retrieval-augmented generation over a restricted document corpus with strict source boundaries, document and PDF data-extraction pipelines that normalize inconsistent real-world specifications, NLP classification pipelines, semantic search, and conversational intake that converts informal user language into exact structured data. The engineer works within a small senior delivery team - a Solution Architect who owns the technical design, and a Senior Full-Stack Developer who consumes the engineer's APIs - and demonstrates completed work in fortnightly sprint reviews with client stakeholders present.
EolasFlow is an AI-native engineering team: AI-assisted development (Claude Code, Cursor, GitHub Copilot or equivalent) is the standard working method, and candidates are expected to already work this way.
Key Responsibilities
- Design and build a production RAG system: chunking and embedding strategy, vector store, retrieval evaluation, citation-grounded answering, and strict source-boundary enforcement with refusal on out-of-bound queries
- Build document-intelligence pipelines: PDF and table extraction from inconsistent source documents, unit and format normalization, deduplication, and human-audit workflows
- Build NLP pipelines for content classification (signal vs noise), entity extraction and enrichment, and automated draft generation matched to a defined editorial voice
- Build semantic search mapping natural-language intent to structured capability data
- Build LLM-guided conversational intake converting informal language into precise structured specifications
- Establish evaluation discipline: evaluation sets and regression harnesses ahead of tuning, evaluations running in CI, quantified quality reporting
- Monitor and optimize cost, latency and quality across all LLM usage; make provider and model trade-offs explicit
- Expose all capabilities as clean, documented APIs for consumption by the application layer
- Present completed work in fortnightly sprint reviews
- Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
Required Skills and Experience
- Minimum 4 years building ML/NLP/LLM systems in production; strong Python (FastAPI or similar for serving)
- Production RAG experience: candidates must be able to walk through a shipped system — architecture, evaluation results, failure modes and remediation
- LLM engineering: prompt design, structured output (JSON schema / function calling), multi-provider model selection (OpenAI, Anthropic, open-weight models),
cost and latency optimization
- Vector stores (pgvector, Qdrant, Pinecone or Weaviate); retrieval evaluation and hallucination control
- Document intelligence: PDF and table extraction from inconsistent real-world documents (Unstructured, Textract, Docling or custom pipelines)
- Evaluation discipline: builds evaluation sets and regression harnesses as standard practice and can quantify quality improvements
- Classic NLP fundamentals beyond prompting classification, named-entity recognition, entity resolution
- Data pipeline orchestration (Airflow, Prefect or similar); compliant API and web data ingestion (rate limiting, terms-of-service awareness)
- Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
- Experience building agentic pipelines (tool use, multi-step agents) in production
- Daily, fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot or equivalent); this will be assessed through a live practical exercise during selection
- Fluent written and spoken English; able to present work to non-technical stakeholders
Desirable
- Small language model (SLM) fine-tuning: LoRA/QLoRA adaptation of open-weight models (Llama, Mistral, Phi or similar class) for classification and style/domain adaptation, including serving and deployment (vLLM, Ollama or similar) valued as a cost- and latency-optimization path for high-volume pipeline tasks
- Hybrid retrieval and re-ranking (BM25 combined with dense retrieval)
- Experience with technical or industrial specification data
- Content personalization or recommender systems
📌 AI Engineer, Exp: 4-7 Yrs, Noida-58, WFO
🏢 Ramp infotech
📍 Noida