19 Sep
|
Hexaware
|
India
Description
Test Automation + Agent/LLM Evaluation (Accuracy, Hallucination, Regression) Role Summary: Responsible for designing, developing, and executing automated testing frameworks for AI/LLM-based applications and intelligent agents. The role focuses on validating model performance, response accuracy, hallucination detection, regression testing, and overall reliability of AI solutions.
Key Responsibilities
- Develop and maintain automated test frameworks for AI agents and LLM-powered applications.
- Evaluate LLM responses for accuracy, relevance, completeness, and consistency.
- Design test cases to identify and measure hallucinations, incorrect outputs, and bias.
- Perform regression testing to ensure model updates do not impact existing functionality.
- Create evaluation datasets, benchmarks, and test scenarios for agent workflows.
- Monitor key AI quality metrics such as precision, recall, response quality, and task success rate.
- Collaborate with AI engineers, data scientists,
and product teams to improve model performance.
- Generate testing reports, insights, and recommendations for continuous improvement.
Required Skills
- Experience in test automation using Python, Selenium, Playwright, PyTest, or similar tools.
- Understanding of Generative AI, LLMs, AI agents, and prompt engineering.
- Knowledge of LLM evaluation frameworks such as LangSmith, DeepEval, Ragas, Promptfoo, or similar.
- Experience with API testing and automated validation pipelines.
- Robust analytical and troubleshooting skills.
Preferred Qualifications
- Experience with OpenAI, Azure OpenAI, Anthropic, Gemini, or other LLM platforms.
- Knowledge of MLOps, AI model monitoring, and CI/CD integration.
- Familiarity with AI governance, responsible AI, and model quality assessment.
📌 Test Lead - API Automation (India)
🏢 Hexaware
📍 India