21 Aug
|
Valiance Solutions
|
Noida
21 Aug
Valiance Solutions
Noida
About the Role
We are looking for an AI Engineer to work on an enterprise AI platform that is transforming how public-sector organizations evaluate tenders and procurement documents.
Our platform uses Generative AI, Agentic AI, document intelligence and open-source foundation models to automate complex document-heavy workflows involved in tender evaluation. The product is already deployed in large Public Sector Undertakings (PSUs) and runs on NVIDIA GPU infrastructure using open-source AI models.
This role sits at the intersection of AI research and production engineering . You will experiment with new models and techniques, but the goal is not research for its own sake — it is to improve accuracy, reduce latency and cost, and make the product increasingly autonomous.
What You Will Work On
- Experiment with and evaluate the latest open-source LLMs and VLMs for document understanding, extraction, reasoning and classification.
- Benchmark models for accuracy, latency, throughput and GPU utilization and make recommendations for production adoption.
- Optimize AI inference pipelines to improve response time, throughput and infrastructure efficiency .
- Develop and deploy Agentic AI workflows that can autonomously execute multi-step document processing and tender evaluation tasks.
- Design agents capable of planning, reasoning, tool use, information retrieval and validation across complex procurement workflows.
- Build production-grade pipelines for processing large, complex and highly unstructured documents .
- Work on techniques such as RAG, structured extraction, reranking, model routing, prompt engineering and evaluation .
- Fine-tune or adapt models where appropriate to improve performance on domain-specific procurement tasks.
- Work closely with product and domain teams to translate business requirements into AI capabilities.
- Profile and optimize GPU workloads across NVIDIA GPU infrastructure .
- Develop robust evaluation frameworks to continuously measure LLM accuracy, hallucination, latency and reliability .
- Stay current with developments in open-source models, inference frameworks and agentic AI, and rapidly prototype promising approaches.
What We’re Looking For
Must Have
- 3–5 years of hands-on experience building and deploying AI/ML systems.
- Solid Python programming and software engineering fundamentals.
- Strong understanding of Generative AI / LLMs and modern AI application architectures.
- Hands-on experience deploying models in production environments.
- Experience with open-source LLMs such as Llama, Qwen, Mistral, DeepSeek or similar models.
- Experience with GPU-based inference and optimization .
- Strong understanding of model performance metrics including latency, throughput, concurrency and GPU utilization .
- Experience with Docker and Linux ; Kubernetes is a strong plus.
- Ability to read research papers, understand new techniques and turn them into working production systems.
- Strong problem-solving mindset and willingness to experiment rapidly.
Good to Have
- Experience with NVIDIA GPUs, CUDA, TensorRT, Triton Inference Server,
vLLM or NVIDIA NIM .
- Experience with Agentic AI / AI agents , tool calling, orchestration and multi-agent workflows.
- Experience with RAG, vector databases, embeddings and reranking .
- Experience with LLM/VLM evaluation and benchmarking .
- Experience with document AI, OCR, document understanding or multimodal models.
- Experience with PyTorch, Hugging Face Transformers and PEFT/LoRA .
- Familiarity with cloud or on-prem GPU deployments.
- Experience working with large-scale enterprise AI deployments.
What Success Looks Like In this role, you will be expected to move beyond simply integrating LLM APIs. You will help us answer questions such as:
- Can we make the model 20% more accurate?
- Can we reduce inference latency by 30%?
- Can we run the workload on fewer GPUs?
- Can an AI agent autonomously complete a document evaluation workflow that currently requires human intervention?
- Can we reliably deploy a new open-source model into production without compromising quality?
You will have significant ownership in taking these ideas from experiment → benchmark → engineering → production . Why Join Us?
You will work on an AI product that is already solving real enterprise problems in production , rather than building prototypes that never leave the lab.
The role offers an opportunity to work across:
Open-source AI models → AI research → GPU optimization → Agentic AI → Production engineering → Enterprise deployment
If you enjoy going deep into models, experimenting with new AI techniques, writing production-grade code and seeing your work directly improve a live AI product, this role is for you.
📌 Generative AI Engineer (Noida)
🏢 Valiance Solutions
📍 Noida