20 Aug
|
Hiringeye Solutions
|
Bengaluru
20 Aug
Hiringeye Solutions
Bengaluru
We are looking for an AI Eval Engineer to build evaluation systems and tools for LLM-powered education products. The role involves Python development, AI workflow testing, model evaluation, data analysis, and improving the quality, accuracy, and safety of AI outputs.
Key Responsibilities
- Build and maintain LLM evaluation pipelines and testing tools.
- Develop Python scripts for batch testing, model comparison, regression testing, and data analysis.
- Create test datasets, golden datasets, rubrics, and evaluation criteria.
- Evaluate AI workflows involving prompts, RAG, structured outputs, tool calling, and multi-step reasoning.
- Identify hallucinations, quality issues, regressions, and recurring AI failure patterns.
- Collaborate with Engineering, Product, QA, and Curriculum teams to improve AI products.
Required Skills
- 3+ years of experience in Software Development, AI/ML, AI Evaluation, QA Automation, or Data Analysis.
- Robust Python and scripting skills.
- Experience with LLM APIs, JSON, datasets, notebooks, and API-based workflows.
- Understanding of AI evaluation, prompt testing, regression testing, and quality measurement.
- Strong analytical and debugging skills.
Preferred Skills
- Experience with LLM-as-a-Judge, RAG, tool calling, and multi-step AI workflows.
- Experience in EdTech, curriculum, tutoring, or educational products.
- Familiarity with dashboards, annotation tools, Git, and experiment tracking.
📌 AI Eval Engineer (Bengaluru)
🏢 Hiringeye Solutions
📍 Bengaluru