Job Summary: We are looking for a GenAI / LLM Agent Test Engineer to assure quality, reliability, and responsible behavior of LLM-based and RAG-based agents, including single-agent and multi-agent workflows. The role focuses on prompt engineering, hallucination detection, agent reasoning validation, and automated evaluation of GenAI outputs using contemporary LLM testing frameworks.
Key Responsibilities: •
Test and validate LLM-based and RAG-based agents, including reasoning, memory, and decision-making behavior •
Design and execute prompt engineering and adversarial testing to uncover hallucinations, bias, and edge cases •
Perform hallucination testing, relevance checks, and factual accuracy validation •
Implement unit tests for LLMs using LLM-as-a-Judge techniques •
Validate agent reasoning chains, memory states, and conversation persistence •
Embed LLMs within testing pipelines to: •
Detect hallucinations • Run adversarial prompt testing
- Perform automated evaluation of agent outputs •
- Ensure AI Ethics and Responsible AI compliance in testing coverage •
- Collaborate with platform, data, and GenAI teams for continuous quality improvement