09 Sep
|
Dentsu Global Services
|
Karnataka
09 Sep
Dentsu Global Services
Karnataka
AI QA Automation Engineer - Immediate Joiners only (0-15 Days) Bangalore / Pune / Mumbai Exp Range - 5-8 years Key Responsibilities Eval Framework & Quality Delivery Design and build eval/golden-set regression framework for non-deterministic GenAI agent output, from scratch. Partner with each AI Engineer to scope and build a golden-set test suite for every agent they ship, landing alongside the agent rather than after. Define and track what "done" looks like for agent output quality, including golden-set pass rate and the rate of defects/false positives that escape to real users. Extend and maintain general automated test coverage across the agents module - backend (xUnit) and frontend E2E (Playwirght) - where gaps exist alongside the eval work. Flag where eval suite depth or scope needs to flex to keep pace with the delivery cadence, and work with the lead engineer to re-balance rather than let a quality backlog build up silently. Beyond the agents module, contribute to and help scale test automation practice across the wider platform - frontend E2E (Playwright) and backend (xUnit) - as the team's senior dedicated test engineer. Required Experience & Skills 3-5 years' experience in a software engineering or test automation role. Experience designing test strategy from scratch, not just executing an existing suite. Experience with C#/.NET testing (xUnit or similar) and/or JavaScript E2E testing frameworks (Playwright or similar).
Comfortable reading and reasoning about LLM/agent output well enough to write meaningful test assertions against it. Ability to work independently across two engineers' agents in parallel, without becoming a bottleneck on their delivery cadence. Experience working in a sprint-based development team. Desirable / Nice-to-Have Skills Experience with some of the following is beneficial but not required: Hands-on experience evaluating GenAI/LLM applications (golden sets, LLM-as-judge, eval frameworks such as Promptfoo, DeepEval, Ragas, or similar). Experience with Google Vertex AI, Gemini, or a comparable LLM platform. Exposure to MCPs, or tool-use integrations for LLMs. Familiarity with ad platform APIs (Google Ads, Microsoft Advertising, Meta, TikTok, DV360). Experience with prompt management tooling (we use PostHog). API testing experience (e.g. via Playwright's API testing capabilities). Tech Workplace The product is built using: Backend : C# (.NET 10), Microsoft Agent Framework, Google Vertex AI (Gemini) Frontend : ReactJS Data : MySQL, Firestore, Google Cloud Storage Infrastructure : Google Cloud Platform (Cloud Run, Cloud Functions, Cloud Scheduler, Docker) Prompt Management : PostHog Testing : xUnit (C#), Playwright (frontend E2E) - general repo conventions. The eval/golden-set framework for agent output does not exist yet and is what this role builds and owns
📌 AI QA Test Engineer (Karnataka)
🏢 Dentsu Global Services
📍 Karnataka