15 Aug
|
Valiance Solutions
|
Noida
15 Aug
Valiance Solutions
Noida
About the Role
We are looking for an AI QA Engineer to help ensure the quality, accuracy, reliability and performance of an enterprise AI platform used by large Public Sector Undertakings (PSUs) for tender and procurement evaluation.
The platform uses Generative AI, Agentic AI, Document AI and open-source models running on NVIDIA GPU infrastructure to automate complex, document-heavy procurement workflows.
This is not a traditional software QA role. You will work at the intersection of AI evaluation, software testing and enterprise product quality, testing systems where outputs can be probabilistic and where correctness depends on context, reasoning and source documents.
What You Will Work On
- Design and execute test strategies for AI-powered document workflows.
- Build test datasets and evaluation benchmarks for tender documents, technical specifications, eligibility criteria, commercial conditions and other procurement content.
- Evaluate AI outputs for accuracy, completeness, consistency, relevance and hallucination.
- Test LLM, VLM, RAG and Agentic AI workflows across different document types and scenarios.
- Validate whether AI-generated answers and extracted information are correctly grounded in source documents.
- Develop automated evaluation pipelines to measure AI quality at scale.
- Perform regression testing whenever models, prompts, agents, retrieval pipelines or application components are changed.
- Test AI agents for correct planning, tool usage, decision-making, workflow execution and failure handling.
- Identify edge cases involving ambiguous, incomplete, conflicting or poorly formatted procurement documents.
- Benchmark different AI models and configurations based on accuracy, latency, reliability and cost.
- Test application APIs, backend services and end-to-end workflows.
- Work with AI engineers to identify root causes of model and pipeline failures.
- Track and analyze production issues and convert them into repeatable test cases.
- Establish quality gates and acceptance criteria for AI models before production deployment.
- Validate performance under realistic concurrency, document volume and workload conditions.
What We're Looking For
Must Have
- 3–5 years of experience in software QA/testing, with strong experience in automation.
- Strong understanding of test case design, test planning, defect management and regression testing.
- Strong Python programming skills and experience building test automation.
- Experience testing APIs, backend systems and web applications.
- Strong analytical and problem-solving skills.
- Ability to work with large and complex datasets and identify subtle quality issues.
- Strong attention to detail and ability to systematically investigate failures.
- Experience working closely with developers and product teams in an Agile environment.
Strongly Preferred
- Hands-on experience testing Generative AI / LLM applications.
- Understanding of LLM evaluation concepts such as hallucination, grounding, relevance, faithfulness and consistency.
- Experience with RAG evaluation and document-based AI systems.
- Familiarity with LLM evaluation frameworks such as RAGAS, DeepEval or similar tools.
- Experience testing AI agents / Agentic AI workflows.
- Familiarity with OCR, Document AI, information extraction and multimodal AI.
- Experience with tools such as PyTest, Selenium, Playwright, Postman or equivalent.
- Understanding of cloud-based AI applications and API-driven architectures.
- Familiarity with LLM/VLM benchmarking and model comparison.
Key Responsibilities
AI Quality & Evaluation
- Define measurable quality metrics for AI features.
- Create gold-standard datasets and expected outputs.
- Evaluate AI responses against ground truth.
- Detect hallucinations,
incorrect reasoning and unsupported conclusions.
- Perform model-to-model and version-to-version comparisons.
Automation
- Build automated test suites for AI and application workflows.
- Automate evaluation of large numbers of documents and AI responses.
- Create regression pipelines that run whenever models or prompts change.
- Integrate automated tests into CI/CD pipelines.
Agent Testing
- Test whether agents correctly interpret instructions and execute multi-step workflows.
- Validate tool calling, retrieval, reasoning and final responses.
- Test agent behavior under failures, missing information and unexpected inputs.
- Identify scenarios where an agent may take an incorrect or unsafe action.
Performance & Reliability
- Test AI systems under realistic production workloads.
- Measure latency, throughput, concurrency and failure rates.
- Identify performance bottlenecks across application and AI layers.
- Validate system behavior across different model and infrastructure configurations.
What Success Looks Like
You will help us move from the AI seems to work to we can quantitatively prove that the AI works.
Success in this role means:
- AI accuracy is measured continuously rather than subjectively.
- New model releases can be evaluated quickly against a standardized benchmark.
- Regressions are detected before reaching customers.
- Hallucinations and document-grounding issues are systematically identified.
- AI agents behave predictably across complex workflows and edge cases.
- Production quality improves with every release.
Why Join Us
- You will be testing real AI systems deployed in production, not demo chatbots.
- You will work closely with AI engineers and product teams on challenging problems involving:
- LLMs → RAG → Document AI → Agentic AI → AI Evaluation → Production QA
- If you are excited about figuring out how to test AI systems that don't always produce the same answer twice, and want to help build reliable enterprise-grade AI, this role is a outstanding opportunity.
📌 Quality Assurance Engineer (Noida)
🏢 Valiance Solutions
📍 Noida