Llm / Agentic Evaluation Rig Engineer Hyderabad

Llm / Agentic Evaluation Rig Engineer Hyderabad

25 Aug
|
Phizenix
|
Hyderabad

25 Aug

Phizenix

Hyderabad

We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes — grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer.

You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar. You are the team's source of truth on whether a model, prompt, or agent change is actually an improvement — and the one who blocks it if it isn't.

What makes this role different
You define "positive enough to ship" — your gates block regressions in grounding and faithfulness from reaching production.
Evidence over vibes — every claim is checked against the verified source data it must be grounded in.
Agentic evaluation — you evaluate multi-step reasoning flows, not just single prompts.
Real leverage — your rig is how the whole AI team moves quick without breaking trust.





Responsibilities

Datasets & Scorers (35%)
Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels
Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity
Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate
Verify generated claims map to verified source data — no unsupported statements

Harnesses & CI Gates (30%)
Build harnesses that run evaluations reproducibly across model, prompt, and agent versions
Wire evaluation into CI so grounding / faithfulness regressions block releases
Track quality over time with dashboards and explicit pass / fail thresholds

Agentic Evaluation (25%)
Evaluate multi-step / agentic flows — routing, tool-use, verification, confirmation
Build trace capture and step-level scoring for agent runs
Detect where a flow silen

📌 Llm / Agentic Evaluation Rig Engineer Hyderabad
🏢 Phizenix
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: llm / agentic evaluation rig engineer hyderabad / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: llm / agentic evaluation rig engineer hyderabad / hyderabad