AI Engineer, Exp: 4-7 Yrs, Noida-58, WFO

AI Engineer, Exp: 4-7 Yrs, Noida-58, WFO

24 Sep
|
Ramp infotech
|
Noida

24 Sep

Ramp infotech

Noida

AI Engineer (LLM / RAG / Document Intelligence)

Role Summary

We are seeking an AI Engineer to build the AI core of a new B2B platform being delivered to a client. Client and project details are confidential and will be disclosed at onboarding under a confidentiality agreement.

The systems operate in a domain where an incorrectly extracted value is a serious defect: accuracy, grounding and controllability take priority over raw capability. The role covers production LLM engineering end-to-end: retrieval-augmented generation over a restricted document corpus with strict source boundaries, document and PDF data-extraction pipelines that normalize inconsistent real-world specifications, NLP classification pipelines, semantic search, and conversational intake that converts informal user language into exact structured data. The engineer works within a small senior delivery team - a Solution Architect who owns the technical design, and a Senior Full-Stack Developer who consumes the engineer's APIs - and demonstrates completed work in fortnightly sprint reviews with client stakeholders present.

EolasFlow is an AI-native engineering team: AI-assisted development (Claude Code, Cursor, GitHub Copilot or equivalent) is the standard working method, and candidates are expected to already work this way.

Key Responsibilities

- Design and build a production RAG system: chunking and embedding strategy, vector store, retrieval evaluation, citation-grounded answering, and strict source-boundary enforcement with refusal on out-of-bound queries
- Build document-intelligence pipelines: PDF and table extraction from inconsistent source documents, unit and format normalization, deduplication, and human-audit workflows




- Build NLP pipelines for content classification (signal vs noise), entity extraction and enrichment, and automated draft generation matched to a defined editorial voice
- Build semantic search mapping natural-language intent to structured capability data
- Build LLM-guided conversational intake converting informal language into precise structured specifications
- Establish evaluation discipline: evaluation sets and regression harnesses ahead of tuning, evaluations running in CI, quantified quality reporting
- Monitor and optimize cost, latency and quality across all LLM usage; make provider and model trade-offs explicit
- Expose all capabilities as clean, documented APIs for consumption by the application layer
- Present completed work in fortnightly sprint reviews
- Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency

Required Skills and Experience

- Minimum 4 years building ML/NLP/LLM systems in production; strong Python (FastAPI or similar for serving)
- Production RAG experience: candidates must be able to walk through a shipped system — architecture, evaluation results, failure modes and remediation
- LLM engineering: prompt design, structured output (JSON schema / function calling), multi-provider model selection (OpenAI, Anthropic, open-weight models),



cost and latency optimization
- Vector stores (pgvector, Qdrant, Pinecone or Weaviate); retrieval evaluation and hallucination control
- Document intelligence: PDF and table extraction from inconsistent real-world documents (Unstructured, Textract, Docling or custom pipelines)
- Evaluation discipline: builds evaluation sets and regression harnesses as standard practice and can quantify quality improvements
- Classic NLP fundamentals beyond prompting classification, named-entity recognition, entity resolution
- Data pipeline orchestration (Airflow, Prefect or similar); compliant API and web data ingestion (rate limiting, terms-of-service awareness)
- Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
- Experience building agentic pipelines (tool use, multi-step agents) in production
- Daily, fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot or equivalent); this will be assessed through a live practical exercise during selection
- Fluent written and spoken English; able to present work to non-technical stakeholders

Desirable

- Small language model (SLM) fine-tuning: LoRA/QLoRA adaptation of open-weight models (Llama, Mistral, Phi or similar class) for classification and style/domain adaptation, including serving and deployment (vLLM, Ollama or similar) valued as a cost- and latency-optimization path for high-volume pipeline tasks
- Hybrid retrieval and re-ranking (BM25 combined with dense retrieval)
- Experience with technical or industrial specification data
- Content personalization or recommender systems

📌 AI Engineer, Exp: 4-7 Yrs, Noida-58, WFO
🏢 Ramp infotech
📍 Noida

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer, exp: 4-7 yrs, noida-58, wfo / noida