30 Sep
|
Snow Planet
|
Hyderabad
30 Sep
Snow Planet
Hyderabad
Job Description
We build the evaluation layer that validates AI agents before they reach customers - an automated system that scores across large volumes of agent traces.
Youll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
What youll do
- Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety - across LLM-as-a-Judge, trajectory-based, and human evaluation
- Take problems from research question to prototype to shipped feature, owning them end to end
- Build and harden the pipelines and scoring logic behind customer-facing evaluation
- Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels Qualifications
What were looking for
- 5+ years in ML, applied AI, prompt engineering, agentic AI,
including shipping something real to users
- Strong Python skills
- Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context
- Experience designing evaluation methodologies, not just running evaluations
- Hands-on production work with LLM APIs - prompt engineering, structured output, cost and latency tradeoffs
- Explicit communication with technical and non-technical audiences
- Good to have Experience with AI-assisted development tools (Claude code, Windsurf, or similar)
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Research Scientist, Agentic AI (Hyderabad)
🏢 Snow Planet
📍 Hyderabad