ABOUT : Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination). KEY RESPONSIBILITIES Build test plans and automation for AI-native application features (functional AI-specific) Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rate Run regression testing across model/prompt/config changes to catch silent quality drift Red-team AI features for edge cases and adversarial inputs where relevant Build automated eval pipelines integrated into CI/CD Partner with AI Architects to define testability requirements before build starts Own the quality gate before any AI feature ships to production Communicate quality risk to delivery leadership in terms they can act on Train delivery teams on AI-specific testing practices Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case Define eval acceptance thresholds per engagement and gate releases on them Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read Instrument production evals and drift monitoring,
feeding failures back into the golden datasets REQUIREMENTS & SKILLS 5–9 yrs QA/test engineering, with 2 yrs testing AI/ML-powered features specifically Solid test automation skills (Python-based frameworks, CI/CD integration) Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them Statistically literate enough to interpret model evaluation metrics, not just pass/fail results Familiarity with red-teaming methodologies for AI systems Clear, assertive communicator — willing to block a release over a quality concern Detail-oriented and methodical under delivery-timeline pressure Collaborative but independent — doesn't rubber-stamp under delivery pressure Explains quality risk in business-impact terms, not just technical jargon Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency Experience wiring evals and drift monitoring into CI/CD and production observability
📌 Forward Deployed Engineer - AI Assurance (India)
🏢 Systems
📍 India