Senior AI Evaluation Engineer / Staff AI Evaluation Engineer
Experience - 6+ Years
Role Overview
We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, quality, and safety evaluation of next-generation Agentic AI applications.
This role sits at the intersection of AI/ML, Software Engineering, Quality Engineering, and Test Automation. The ideal candidate should have solid hands-on experience building automated evaluation frameworks and testing LLM, RAG, LangGraph, LangChain, and multi-agent applications.
The candidate will work closely with Engineering, Product, Data, AI/ML, and Quality Engineering teams to establish measurable AI quality standards and build automated regression and evaluation suites for production AI systems.
Key Responsibilities
- Design and implement automated evaluation frameworks for LLM, RAG, and Agentic AI applications.
- Develop measurable evaluation criteria for AI quality, accuracy, relevance, reliability, safety, and consistency.
- Build and maintain golden datasets, benchmark datasets, test datasets, and evaluation suites.
- Perform functional, regression, integration, end-to-end, adversarial, and safety testing of AI applications.
- Evaluate LangGraph-based and multi-agent applications at the node, state-transition, routing, and workflow levels.
- Build automated evaluation workflows using Python and modern testing frameworks.
- Use LangSmith for tracing, debugging, dataset management, experiments, and AI evaluations.
- Design and execute adversarial testing and identify failure modes, hallucinations, prompt vulnerabilities, and unexpected agent behavior.
- Evaluate AI applications for Responsible AI, security, privacy, and safety considerations.
- Work with LangChain and LangGraph,
including sub-graphs and conditional routing.
- Integrate AI evaluation and automated testing into GitHub-based CI/CD workflows.
- Work with AWS and Amazon Bedrock for AI application testing and evaluation.
- Analyze LLM outputs, embeddings, vector search, prompts, RAG pipelines, and agent behavior.
- Develop automated regression suites to detect model, prompt, retrieval, and application-level quality degradation.
- Collaborate with engineering teams to troubleshoot failures and improve AI system reliability.
- Document evaluation methodologies, test results, quality metrics, and identified risks.
- Independently investigate ambiguous technical problems and drive them toward measurable solutions.
Required Qualifications
- 6+ years of experience in Software Engineering, Quality Engineering, Test Automation, AI Engineering, Machine Learning, or a related technical field.
- Hands-on experience testing or evaluating LLM, NLP, Machine Learning, or AI applications.
- Strong hands-on Python programming and automation experience.
- Experience creating and maintaining golden datasets, benchmark datasets, or test datasets.
- Hands-on experience with adversarial testing.
- Knowledge of AI Safety, Responsible AI, Security, and Privacy evaluation.
- Strong structural understanding of LangChain, LangGraph, and LangSmith.
- Hands-on experience with LangSmith for tracing, debugging, experiments, datasets, and evaluations.
- Strong GitHub experience including repositories, branching, pull requests, code reviews, and CI/CD workflows.
- Experience with Claude Code or similar AI-assisted development tools.
- Experience with AWS and Amazon Bedrock.
- Strong knowledge of REST APIs, JSON, SQL, and modern application architectures.
- Strong understanding of LLMs, prompt engineering, embeddings, vector search, RAG, and AI agents.
- Experience with automated, regression, integration, and end-to-end testing.
- Hands-on experience evaluating LangGraph-based or multi-agent applications at node and state-transition levels.
- Solid analytical, troubleshooting, communication, and problem-solving skills.
- Ability to work independently and take ownership of complex and ambiguous technical challenges.
Primary Mandatory Skills:
- 6+ years of relevant Software/QA/AI Engineering experience
- Python
- LLM / GenAI Application Testing & Evaluation
- LangChain
- LangGraph
- LangSmithAI/LLM Evaluation Frameworks
- Golden / Benchmark Dataset Creation
- Adversarial Testing
- RAG
- Prompt Engineering
- AI Agents / Multi-Agent Systems
- Automated Testing & Regression Testing
Secondary Mandatory Skills:
- AWS
- Amazon Bedrock
- GitHub & CI/CD
- REST APIs
- JSON
- SQL
- AI Safety / Responsible AI
- Security & Privacy Evaluation
- Embeddings & Vector Search
- Integration & End-to-End Testing
- Claude Code or similar AI-assisted development tools
Preferred Candidate Profile
Candidates with hands-on experience in AI evaluation engineering, GenAI quality engineering, LLM testing, Agentic AI testing, LangGraph evaluation, LangSmith, RAG evaluation, and automated AI regression frameworks will be particularly relevant for this position.
📌 Senior AI Evaluation Engineer | Remote | (India)
🏢 Haparz
📍 India