01 Sep
|
SIDE
|
Hyderabad
Role Details
Contractual Role - 3 months (will be extended depending on performance and project requirement)
Location - Hyderabad (Nanakramguda)
Key Responsibilities
- Test and validate AI-generated insights, recommendations, and decision-making workflows.
- Evaluate LLM and RAG systems for accuracy, relevance, consistency, factuality, and hallucinations.
- Validate retrieval quality, context relevance, grounding, and response quality in RAG systems.
- Test AI agents and autonomous workflows across functional, negative, and edge-case scenarios.
- Define AI evaluation criteria, test datasets, quality metrics, and validation processes.
- Perform regression testing for models, prompts, RAG configurations, and AI workflows.
- Collaborate with AI/ML engineers to identify issues and improve AI system quality
Required Skills
- Strong understanding of AI/ML and Generative AI testing.
- 5-9 years of overall experience with GenAI experience of atleast 2-3+ years in LLM/GenAI testing
- Hands-on experience testing LLM and RAG-based applications.
- Hands-on LLM testing (GPT, Claude, Gemini, Llama)
- Experience testing at least one RAG-based application
- Python automation, PyTest, API testing skills
- Hallucination, relevance, groundedness and response quality validation
- Exposure to evaluation frameworks such as Ragas, DeepEval, Promptfoo, LangSmith or TruLens
- Robust analytical and problem-solving skills.
📌 AI Quality Engineer | Contractual | Hyderabad
🏢 SIDE
📍 Hyderabad