Senior AI Evaluation Engineer | Remote | (India)

Senior AI Evaluation Engineer | Remote | (India)

28 Sep
|
Haparz
|
India

28 Sep

Haparz

India

Senior AI Evaluation Engineer / Staff AI Evaluation Engineer

Experience - 6+ Years

Role Overview

We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, quality, and safety evaluation of next-generation Agentic AI applications.

This role sits at the intersection of AI/ML, Software Engineering, Quality Engineering, and Test Automation. The ideal candidate should have solid hands-on experience building automated evaluation frameworks and testing LLM, RAG, LangGraph, LangChain, and multi-agent applications.

The candidate will work closely with Engineering, Product, Data, AI/ML, and Quality Engineering teams to establish measurable AI quality standards and build automated regression and evaluation suites for production AI systems.

Key Responsibilities

- Design and implement automated evaluation frameworks for LLM, RAG, and Agentic AI applications.
- Develop measurable evaluation criteria for AI quality, accuracy, relevance, reliability, safety, and consistency.
- Build and maintain golden datasets, benchmark datasets, test datasets, and evaluation suites.
- Perform functional, regression, integration, end-to-end, adversarial, and safety testing of AI applications.
- Evaluate LangGraph-based and multi-agent applications at the node, state-transition, routing, and workflow levels.
- Build automated evaluation workflows using Python and modern testing frameworks.
- Use LangSmith for tracing, debugging, dataset management, experiments, and AI evaluations.
- Design and execute adversarial testing and identify failure modes, hallucinations, prompt vulnerabilities, and unexpected agent behavior.
- Evaluate AI applications for Responsible AI, security, privacy, and safety considerations.
- Work with LangChain and LangGraph,



including sub-graphs and conditional routing.
- Integrate AI evaluation and automated testing into GitHub-based CI/CD workflows.
- Work with AWS and Amazon Bedrock for AI application testing and evaluation.
- Analyze LLM outputs, embeddings, vector search, prompts, RAG pipelines, and agent behavior.
- Develop automated regression suites to detect model, prompt, retrieval, and application-level quality degradation.
- Collaborate with engineering teams to troubleshoot failures and improve AI system reliability.
- Document evaluation methodologies, test results, quality metrics, and identified risks.
- Independently investigate ambiguous technical problems and drive them toward measurable solutions.

Required Qualifications

- 6+ years of experience in Software Engineering, Quality Engineering, Test Automation, AI Engineering, Machine Learning, or a related technical field.
- Hands-on experience testing or evaluating LLM, NLP, Machine Learning, or AI applications.
- Strong hands-on Python programming and automation experience.
- Experience creating and maintaining golden datasets, benchmark datasets, or test datasets.
- Hands-on experience with adversarial testing.
- Knowledge of AI Safety, Responsible AI, Security, and Privacy evaluation.
- Strong structural understanding of LangChain, LangGraph, and LangSmith.
- Hands-on experience with LangSmith for tracing, debugging, experiments, datasets, and evaluations.




- Strong GitHub experience including repositories, branching, pull requests, code reviews, and CI/CD workflows.
- Experience with Claude Code or similar AI-assisted development tools.
- Experience with AWS and Amazon Bedrock.
- Strong knowledge of REST APIs, JSON, SQL, and modern application architectures.
- Strong understanding of LLMs, prompt engineering, embeddings, vector search, RAG, and AI agents.
- Experience with automated, regression, integration, and end-to-end testing.
- Hands-on experience evaluating LangGraph-based or multi-agent applications at node and state-transition levels.
- Solid analytical, troubleshooting, communication, and problem-solving skills.
- Ability to work independently and take ownership of complex and ambiguous technical challenges.

Primary Mandatory Skills:

- 6+ years of relevant Software/QA/AI Engineering experience
- Python
- LLM / GenAI Application Testing & Evaluation
- LangChain
- LangGraph
- LangSmithAI/LLM Evaluation Frameworks
- Golden / Benchmark Dataset Creation
- Adversarial Testing
- RAG
- Prompt Engineering
- AI Agents / Multi-Agent Systems
- Automated Testing & Regression Testing

Secondary Mandatory Skills:

- AWS
- Amazon Bedrock
- GitHub & CI/CD
- REST APIs
- JSON
- SQL
- AI Safety / Responsible AI
- Security & Privacy Evaluation
- Embeddings & Vector Search
- Integration & End-to-End Testing
- Claude Code or similar AI-assisted development tools

Preferred Candidate Profile

Candidates with hands-on experience in AI evaluation engineering, GenAI quality engineering, LLM testing, Agentic AI testing, LangGraph evaluation, LangSmith, RAG evaluation, and automated AI regression frameworks will be particularly relevant for this position.

📌 Senior AI Evaluation Engineer | Remote | (India)
🏢 Haparz
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai evaluation engineer | remote | (india) / india