The role
You make sure our AI features actually work, and stay working. AI systems are non-deterministic, so classic QA is not enough. You build the evals, checks and observability that tell us when a model or agent has quietly gotten worse.
You have tested AI or model-backed features before and understand why the same input can pass today and fail tomorrow.
What you will do
Build test suites and eval sets for model and agent outputs, including regression checks when models or prompts change.
Set up AI observability: tracing prompts, responses, tool calls, latency, cost and failure rates.
Define what valuable output looks like for each feature and build automated ways to score it.
Catch quality drift, hallucinations, unsafe output and broken tool calls before users do.
Feed findings back to the engineers building the models and workflows, and flag functional breakage clearly.
What we are looking for
Hands-on with AI-driven testing and model evaluation is essential: building eval sets, model and prompt evals, RAG testing, and LLM-as-a-Judge scoring. These are core to the role, not nice-to-haves.
One to two years in QA, testing or a quality role, including hands-on work with AI or model-backed features.
You understand why AI systems need evals, not just pass or fail test cases.
Comfortable writing scripts in Python or JavaScript to automate checks.
You can read logs and traces and reason about what a model or agent did.
📌 AI QA and Observability Engineer (Noida)
🏢 Hungama
📍 Noida