AI Quality Analyst (India)

AI Quality Analyst (India)

06 Sep
|
Avenue Code
|
India

06 Sep

Avenue Code

India

About The Opportunity

Avenue Code is looking for an AI Quality Analyst to work with Fanatics as part of its AI/ML Platform team.

The AI/ML Platform team builds and supports Fanatics’ internal AI platform, including the gateway, agent runtime, registry, retrieval, and evaluation infrastructure used by teams developing solutions with LLMs and AI agents.

As the team’s dedicated AI Quality Analyst, you will be responsible for evaluating whether AI agents are performing as expected. You will run evaluation sets, review outputs against established rubrics, maintain high-quality ground-truth datasets, identify failure patterns, and provide actionable findings to the engineering and product teams responsible for each agent.

This is a 12-month contract opportunity based in Hyderabad, India, following a hybrid work model.

Responsibilities

- Run evaluation sets against production and pre-release AI agents on a regular cadence and following meaningful changes to prompts, models, tools, or retrieval systems.
- Analyze and report evaluation results, including pass rates, regressions, trends, and recurring failure patterns.
- Review AI agent outputs against defined quality rubrics covering correctness, grounding, tool-use accuracy, safety, tone, and output format.
- Score outputs consistently, document edge cases, and identify areas where evaluation rubrics need clarification or improvement.
- Curate and maintain ground-truth datasets used to evaluate AI and LLM-based applications.
- Create, review, and validate golden answers and expected outputs for evaluation cases.
- Classify failures by root cause, including prompt issues, retrieval failures, incorrect tool usage, or model-related errors.
- Identify and retire outdated, ambiguous, or low-quality evaluation cases.




- Compare LLM-as-a-judge evaluation results against human review and identify discrepancies between automated and human scoring.
- Provide agent owner teams with reproducible examples, severity assessments, and suspected root causes for identified issues.
- Partner with engineering and product teams to track fixes and re-run evaluations until issues are resolved.
- Convert production failures, user feedback, and user corrections into new evaluation cases to continuously improve evaluation coverage.
- Use Python and SQL to execute evaluation scripts, analyze results, investigate traces, and identify patterns in agent behavior.
- Write explicit and concise quality reports, bug reports, and evaluation summaries for technical and product stakeholders.

Required Qualifications

- 3+ years of professional experience in Quality Assurance, Data Annotation, Data Analytics, AI Quality, or another evidence-driven role.
- At least 1 year of hands-on experience reviewing or evaluating outputs from LLMs, conversational AI systems, chatbots, or AI agents.
- Strong ability to evaluate open-ended AI-generated content consistently using predefined scoring rubrics.
- Ability to clearly explain and defend evaluation decisions based on defined criteria and evidence.
- Working knowledge of Python and SQL for running evaluation scripts, analyzing datasets, and investigating results.
- Understanding of common failure modes in LLM applications,



Retrieval-Augmented Generation (RAG), tool calling, and agent workflows.
- Ability to distinguish between issues caused by retrieval, prompting, tools, data, and model hallucinations.
- Experience creating or maintaining ground-truth datasets, golden answers, or labeled evaluation datasets.
- Strong analytical skills and attention to detail.
- Excellent written and verbal communication skills, with the ability to document findings clearly for engineers and product stakeholders.
- Fluent written and spoken English.
- Ability to maintain regular overlap with US working hours for team collaboration and pod syncs.
- Familiarity with AI observability, tracing, or evaluation platforms such as Langfuse, Arize Phoenix, or similar tools is a plus.

Nice To Have Skills

- Familiarity with Langfuse, Arize Phoenix, or similar tracing and eval tools
- Understanding of RAG
- Understanding of tool calling
- Understanding of agent loops/workflows
- Ability to distinguish retrieval misses vs. model hallucinations
- Familiarity with LLM-as-a-judge calibration against human review
- Experience curating ground-truth datasets / golden answers

Avenue Code reinforces its commitment to privacy and to all the principles guaranteed by the most accurate global data protection laws, such as GDPR, LGPD, CCPA, and CPRA. Candidate data shared with Avenue Code will be kept confidential and will not be transmitted to disinterested third parties, nor will it be used for purposes other than the application for open positions. As a consultancy company, Avenue Code may share your information with its clients and other companies from the Compass UOL Group to which Avenue Code consultants are allocated to perform their services.

📌 AI Quality Analyst (India)
🏢 Avenue Code
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai quality analyst (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: ai quality analyst (india) / india