21 Aug
|
Hiringeye Solutions
|
Bengaluru
21 Aug
Hiringeye Solutions
Bengaluru
We are looking for an AI Eval Engineer to build evaluation systems and tools for LLM-powered education products. The role involves Python development, AI workflow testing, model evaluation, data analysis, and improving the quality, accuracy, and safety of AI outputs.
Key Responsibilities
Build and maintain LLM evaluation pipelines and testing tools.
Develop Python scripts for batch testing, model comparison, regression testing, and data analysis.
Create test datasets, golden datasets, rubrics, and evaluation criteria.
Evaluate AI workflows involving prompts, RAG, structured outputs, tool calling, and multi-step reasoning.
Identify hallucinations, quality issues, regressions, and recurring AI failure patterns.
Collaborate with Engineering, Product, QA, and Curriculum teams to improve AI products.
Required Skills
3+ years of experience in Software Development, AI/ML, AI Evaluation, QA Automation, or Data Analysis.
Robust Python and scripting skills.
Experience with LLM APIs, JSON, datasets, notebooks, and API-based workflows.
Understanding of AI evaluation, prompt testing, regression testing, and quality measurement.
Solid analytical and debugging skills.
Preferred Skills
Experience with LLM-as-a-Judge, RAG, tool calling, and multi-step AI workflows.
Experience in EdTech, curriculum, tutoring, or educational products.
Familiarity with dashboards, annotation tools, Git, and experiment tracking.
📌 Ai Eval Engineer Bengaluru
🏢 Hiringeye Solutions
📍 Bengaluru