Serve as the technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead s authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.
SECTION B: KEY RESPONSIBILITIES AND RESULTS
Indicate key responsibilities and performance indicators of this role.
For existing role, please indicate additional responsibilities in bold.
1. Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes. 2. Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings. 3. Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team. 4. Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration). 5. Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces. 6. Mentor members of the team when needed and review their work for quality and consistency.
SECTION C:
QUALIFICATIONS / EXPERIENCE / KNOWLEDGE REQUIRED Indicate key knowledge and skills required for this role to perform the tasks to a satisfactory level. To also specify a suitable level of qualification required (i.e. basic, advanced, or professional), where applicable.
Category Essential for this role Good to have Education and Qualifications Bachelor s or Master s degree in Computer Science or a related field Work Experience
6+ years in ML/data/software with strong evaluation or quality focus
- Hands-on experience evaluating LLM or ML systems
Experience with agentic systems or RAG pipelines
Telco AI domain
Technical / Skilled Skills
Please provide at least 3
LLM/agent evaluation design and statistical rigour
Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
- Management presentation Other Task-Specific Knowledge Understanding of telco customer intents and journeys - Responsible-AI and safety evaluation practices
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Agent Evaluation & Instrumentation Engineer (Pune)
🏢 NCS Group
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.