Location: Pune | Experience: 8–11 Years | CTC: Up to ₹33 LPA(576)
Notice Period: Immediate–45 Days | Preference: Local Pune Candidates
Interview: 2 Technical Rounds + Client Round + HR
? MANDATE CRITERIA – NON-NEGOTIABLE
- 6+ years in ML/Data/Software Engineering with strong evaluation/quality focus
- Hands-on LLM/ML evaluation experience
- Robust Python skills
- Experience with Promptfoo / DeepEval / custom evaluation frameworks
- Strong understanding of LLM/Agent evaluation design & statistical rigour
- Hands-on data analysis & metric interpretation
- Experience with tracing & observability tools
- Strong technical documentation, analytical & stakeholder-management skills
- Understanding of Telco customer intents & customer journeys
- 2+ years stability in an organization
? Key Responsibilities
- Design and maintain LLM/Agent evaluation suites, golden sets, regression packs & adversarial tests
- Build continuous evaluation & scoring pipelines
- Conduct quality/operability gate reviews and reproduce evaluation results
- Analyse model regression,
failure modes, defects & evaluation trends
- Calibrate quality thresholds, judges & evaluation datasets
- Monitor evaluation drift and maintain hold-out/golden sets
- Prepare technical quality reports & evaluation findings
- Collaborate with AI/ML, Engineering & Operations teams on quality improvements
- Mentor team members and communicate findings to technical/client stakeholders
⭐ GOOD TO HAVE Agentic AI | RAG | Telco AI | AI Red-Teaming | Responsible AI & Safety | Judge Calibration | Management Presentations Ideal Profile: AI/ML Evaluation + Python + Data Analysis + LLM/Agent Quality with robust communication and client-facing skills.