09 Aug
|
Antal International
|
India
09 Aug
Antal International
India
AI Engineer, Test Generation and Scoring Platform Location: Remote Experience: 2 t0 8 years NP- Immediate We build a platform that tests AI agents before they reach production. You will own the parts that make that possible: the pipelines that generate test cases from an agent definition, the scoring layer that decides whether a response was actually correct, the prompting that drives both, and the backend that runs it all at scale. Why this role exists Testing an agent is not like testing software. The same input does not produce the same output twice. A response can be fluent, confident and wrong. Correctness depends on whether the right tool was called, whether the answer was grounded in the right source, and whether the agent refused when it should have. None of that is captured by a pass or fail assertion. Our platform solves that by generating coverage from the agent's own definition and scoring behaviour along dimensions a customer actually cares about. The engineer who built much of that is moving into a customer facing deployment, and we are hiring their replacement on the core team. You are inheriting real systems with real users, not a greenfield brief. What you will own - Test generation pipelines. Turn an agent definition, a specification or a set of production traces into executable test cases, including negative paths, edge cases and adversarial variants. Make the generated coverage good enough that an engineer does not want to rewrite it by hand. - Scoring and evaluation. Build the layer that decides whether a response was correct. Task completion, tool selection accuracy,
grounding fidelity, refusal correctness, latency and cost. Calibrate scorers against human judgement and keep them stable as models change underneath you. - Prompting as production code. Own the prompts that drive generation and scoring. Version them, test them, and catch the regression when a prompt change quietly degrades output quality across every customer. - Backend services. APIs, job orchestration, queueing, data modelling and storage for test runs at volume. Runs need to be reproducible, resumable and fast enough that nobody avoids triggering them. - Reference agent builds. Build agents on the managed platforms we support, so the product is designed against real platform behaviour rather than documentation. - Reporting. Turn raw run output into something a customer's architect or delivery owner can act on directly, with per scenario traces, failure clustering and a transparent release readiness signal. Must have - Two or more years building production software, with meaningful time on systems backed by large language models. Internships count if the work shipped and had users. - Strong Python. You write code others maintain.
Comfort in TypeScript or a second language is useful but Python is where this work lives. - You have shipped an LLM feature to production. Prompt design, structured output, retries and fallbacks, handling the case where the model returns something unexpected, and keeping cost and latency inside a budget. - Backend fundamentals. REST APIs, asynchronous job processing, data modelling, and deploying to a cloud environment without needing someone else to do it for you. - You have built at least one agent on a managed platform. Any of them. What matters is that you have wired up tool calling and knowledge grounding yourself and seen where it breaks. - Comfort with non determinism. You have opinions about how to test a system that does not return the same answer twice, and you can explain why a simple string match is not the answer. - You ship without being managed. Git, code review, CI, and the judgement to know when something is good enough to release and when it is not. Strongly preferred - Built test infrastructure, scoring harnesses or quality tooling before, in any domain. - Worked across more than one agent or model provider and can speak to how they differ in practice. - Retrieval and grounding systems: vector stores, chunking strategy, citation correctness. - Model finetuning or preference tuning exposure, even at a small scale. - Workflow automation and event driven architecture. - Voice or conversational systems. - Early stage experience, where the roadmap changes and you are close to the customer.
📌 AI Engineer, Test Generation and Scoring Platform (India)
🏢 Antal International
📍 India