Eval Ownership: 6+ years software/ML engineering with demonstrated ownership of an LLM/agent evaluation program that gated real releases. Robust Python (pandas, SQL, pytest). Agent (not just LLM) Evals: Hands-on with trajectory/trace-based evaluation, tool-calling and planning metrics, and pass^k multi-run methodology , output-only eval experience is not enough. Judge Validation: Proven ability to build and statistically validate LLM-as-judge / agent-as-judge evaluators and to articulate why naive judges fail on multi-turn traces and how to mitigate it. Tooling: Practical experience across the current stack , e.g., Braintrust (CI regression), Arize Phoenix (OpenTelemetry-native observability), Promptfoo (OWASP red-team), Galileo, DeepEval/Confident AI, Ragas, Langfuse/LangSmith , and the judgment to know what each metric proves. Trace Fluency:
Comfort reading OpenTelemetry-style agent traces (spans for retrieval, tool calls, sub-agent handoffs) and attributing failures to components. Skeptical Communication: Defends uncomfortable numbers to delivery leadership and client stakeholders; influences engineering without authority.
Desirable Domain: Commercial Real Estate (lease accounting, CAM reconciliation) or document-extraction domains where citation traceability matters. Security Evals: AppSec / red-team background applied to LLM systems. Research Currency: Eval literature (agent-as-a-judge, standardized agent-evaluation work, stochastic-eval research) and integrates it. Compliance: Producing evaluation evidence under SOC 2 / GDPR / Section 508.
📌 Senior AI Evaluation Engineer (Maharashtra)
🏢 Pyramid IT Consulting
📍 Maharashtra
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.