16 Aug
|
Valiance Solutions
|
Noida
16 Aug
Valiance Solutions
Noida
About the Role
We are looking for an AI Engineer to work on an enterprise AI platform that is transforming how public-sector organizations evaluate tenders and procurement documents.
Our platform uses Generative AI, Agentic AI, document intelligence and open-source foundation models to automate complex document-heavy workflows involved in tender evaluation. The product is already deployed in large Public Sector Undertakings (PSUs) and runs on NVIDIA GPU infrastructure using open-source AI models.
This role sits at the intersection of AI research and production engineering. You will experiment with recent models and techniques, but the goal is not research for its own sake — it is to improve accuracy, reduce latency and cost, and make the product increasingly autonomous.
What You Will Work On
- Experiment with and evaluate the latest open-source LLMs and VLMs for document understanding, extraction, reasoning and classification.
- Benchmark models for accuracy, latency, throughput and GPU utilization and make recommendations for production adoption.
- Optimize AI inference pipelines to improve response time, throughput and infrastructure efficiency.
- Develop and deploy Agentic AI workflows that can autonomously execute multi-step document processing and tender evaluation tasks.
- Design agents capable of planning, reasoning, tool use, information retrieval and validation across complex procurement workflows.
- Build production-grade pipelines for processing large, complex and highly unstructured documents.
- Work on techniques such as RAG, structured extraction, reranking, model routing, prompt engineering and evaluation.
- Fine-tune or adapt models where appropriate to improve performance on domain-specific procurement tasks.
- Work closely with product and domain teams to translate business requirements into AI capabilities.
- Profile and optimize GPU workloads across NVIDIA GPU infrastructure.
- Develop robust evaluation frameworks to continuously measure LLM accuracy, hallucination, latency and reliability.
- Stay current with developments in open-source models, inference frameworks and agentic AI, and rapidly prototype promising approaches.
What We’re Looking For
Must Have
- 3–5 years of hands-on experience building and deploying AI/ML systems.
- Strong Python programming and software engineering fundamentals.
- Strong understanding of Generative AI / LLMs and modern AI application architectures.
- Hands-on experience deploying models in production environments.
- Experience with open-source LLMs such as Llama, Qwen, Mistral, DeepSeek or similar models.
- Experience with GPU-based inference and optimization.
- Strong understanding of model performance metrics including latency, throughput, concurrency and GPU utilization.
- Experience with Docker and Linux; Kubernetes is a strong plus.
- Ability to read research papers, understand new techniques and turn them into working production systems.
- Strong problem-solving mindset and willingness to experiment rapidly.
Good to Have
- Experience with NVIDIA GPUs, CUDA, TensorRT, Triton Inference Server,
vLLM or NVIDIA NIM.
- Experience with Agentic AI / AI agents, tool calling, orchestration and multi-agent workflows.
- Experience with RAG, vector databases, embeddings and reranking.
- Experience with LLM/VLM evaluation and benchmarking.
- Experience with document AI, OCR, document understanding or multimodal models.
- Experience with PyTorch, Hugging Face Transformers and PEFT/LoRA.
- Familiarity with cloud or on-prem GPU deployments.
- Experience working with large-scale enterprise AI deployments.
What Success Looks Like In this role, you will be expected to move beyond simply integrating LLM APIs. You will help us answer questions such as:
- Can we make the model 20% more accurate?
- Can we reduce inference latency by 30%?
- Can we run the workload on fewer GPUs?
- Can an AI agent autonomously complete a document evaluation workflow that currently requires human intervention?
- Can we reliably deploy a new open-source model into production without compromising quality?
You will have significant ownership in taking these ideas from experiment → benchmark → engineering → production. Why Join Us?
You will work on an AI product that is already solving real enterprise problems in production, rather than building prototypes that never leave the lab.
The role offers an opportunity to work across:
Open-source AI models → AI research → GPU optimization → Agentic AI → Production engineering → Enterprise deployment
If you enjoy going deep into models, experimenting with new AI techniques, writing production-grade code and seeing your work directly improve a live AI product, this role is for you.
📌 Generative AI Engineer (Noida)
🏢 Valiance Solutions
📍 Noida