AI Quality Assurance Engineer (Bengaluru)

AI Quality Assurance Engineer (Bengaluru)

02 Sep
|
ThirdEye Data
|
Bengaluru

02 Sep

ThirdEye Data

Bengaluru

AI Quality Assurance Engineer

Agent Evaluation

Location

Bengaluru, India (Onsite / Hybrid)

Department

AI Engineering / Quality

Reports to

Head of AI Engineering / Director of QA

Employment type

Full-time

About the Role

You will own the evaluation process for our AI agent system. The system includes models, prompts, tools, memory, and orchestration logic. You will test all of these parts.

You will build the tests that decide if an agent is ready to ship. You will find regressions before customers find them. You will measure agent quality with numbers, not opinions.

What You Will Do

Own the evaluation process

- Build and maintain evaluation suites for the full agent stack.
- Test prompt behavior, tool selection, tool calls, multi-step plans, memory use, error recovery, and task results.
- Define quality metrics for each capability. Examples: task success rate, trajectory correctness, tool-call accuracy, latency, and cost.
- Build and maintain golden datasets, regression suites, and adversarial test sets. Put these datasets under version control.

Build evaluation infrastructure
- Build automated evaluation pipelines. Connect them to CI/CD.
- Make sure each change to a model, prompt, tool, or the harness must pass evaluation before release.
- Use more than one grading method: rule-based checks, exact match, semantic match, human review, and LLM-as-judge.
- Calibrate LLM judges against human labels. Monitor the judges for drift.
- Add tracing and logs to agent runs. Record each step, each tool call, and each cost. Make failures easy to diagnose.

Establish and standardize KPIs and metrics
- Define the KPI set for agent quality. Include accuracy, faithfulness, answer relevance, context precision, context recall, tool-call correctness, hallucination rate, latency, and cost.
- Write standard definitions for each metric. Write standard rubrics and report formats.
- Make sure teams can compare results across models, prompts, and releases.
- Extend evaluation to more than one modality: text, image, audio, and structured documents. Select the correct grading method for each modality.

Analyze failures and drive fixes




- Do structured error analysis. Group failures by type.
- Find the source of each failure: the model, the harness, or a tool.
- Set sample sizes, pass thresholds, and confidence bounds. Agent runs are not deterministic. Plan for variance.
- Work with engineers to reproduce failures and to verify fixes.

Protect the evaluation results
- Find and prevent test contamination, overfitting to benchmarks, and metric gaming.
- Keep a mix of offline evaluations, staged evaluations, and production monitoring.
- Make sure evaluation results predict production behavior.

Set the quality bar
- Define release criteria for agent changes. Approve or block releases with data.
- Write transparent documentation for the evaluation methods.
- Report results to engineering and to leadership.

What We Require AI and agent evaluation (4+ years)
- 4 or more years of hands-on work in evaluation of LLM systems and agents.
- Hands-on skill with RAGAs, Promptfoo, and DeepEval. Skill with Braintrust, LangSmith, W&B; Weave, or OpenAI Evals is a plus.
- Proof that you can create KPIs and metrics from zero: select the metrics, set the baselines, set the thresholds, and get agreement from stakeholders.
- Proof that you can standardize evaluation across teams: shared metric definitions, reusable templates, versioned datasets, and one report format.
- Experience with evaluation in more than one modality: text plus image, audio, video, or structured documents.
- Knowledge of the agent evaluation problem space: trajectory evaluation, outcome evaluation, tool-use correctness, multi-turn behavior, non-determinism, and prompt sensitivity.
- Experience with LLM-as-judge design: rubric design, calibration against human labels, and bias control.




- Knowledge of statistics for evaluation: sampling, significance tests, inter-rater agreement, and variance across runs.
- Strong Python skills. Experience with test harnesses, data pipelines, or evaluation tools that you built.
- Skill with tracing and observability for agent systems. Example: OpenTelemetry-style traces and structured logs of tool calls.

Core software QA
- Knowledge of QA fundamentals: test plans, test-case design, functional tests, regression tests, integration tests, and defect management.
- Hands-on skill with test automation frameworks. Examples: pytest, Playwright, Selenium.
- Hands-on skill with API tests. Examples: Postman, REST-assured.
- Experience with test suites in CI/CD pipelines. Examples: GitHub Actions, Jenkins, GitLab CI.
- Skill with bug tracking and test management tools. Examples: Jira, TestRail, Xray.
- Ability to write clear defect reports that engineers can reproduce.
- Knowledge of non-functional tests: performance, load, and reliability.

What Is a Plus
- Experience with red-teaming, safety evaluations, or adversarial tests of AI systems.
- Experience with human annotation work: labeling guidelines, calibration sessions, and inter-annotator agreement.
- Background in CI/CD, infrastructure-as-code, or platform engineering.
- Contributions to open-source evaluation tools. Published work on evaluation methods.

What Success Looks Like
- 90 days: You know the current agent harness. You found the gaps in the current evaluations. A baseline regression suite runs on each change.
- 6 months: Evaluation results are trusted release gates. Failure types are tracked and decrease. Engineers come to you before they ship.
- 12 months: Other teams use the evaluation platform on their own. Offline results predict production quality. Our quality bar is a competitive advantage.

Why This Role Matters Agents fail in ways that normal software does not. They fail silently. They fail at random. They fail on step ten of a long task. Users trust an AI product only when the evaluation is rigorous. You own that rigor.

📌 AI Quality Assurance Engineer (Bengaluru)
🏢 ThirdEye Data
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai quality assurance engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: ai quality assurance engineer (bengaluru) / bengaluru