03 Sep
|
Dentsu Global Services
|
Bengaluru
03 Sep
Dentsu Global Services
Bengaluru
AI QA Automation Engineer - Immediate Joiners only (0-15 Days)
Bangalore / Pune / Mumbai
Exp Range - 5-8 years
Key Responsibilities
Eval Framework & Quality Delivery
- Design and build --- eval/golden-set regression framework for non-deterministic GenAI agent output, from scratch.
- Partner with each AI Engineer to scope and build a golden-set test suite for every agent they ship, landing alongside the agent rather than after.
- Define and track what "done" looks like for agent output quality, including golden-set pass rate and the rate of defects/false positives that escape to real users.
- Extend and maintain general automated test coverage across the agents module - backend (xUnit) and frontend E2E (Playwirght) - where gaps exist alongside the eval work.
- Flag where eval suite depth or scope needs to flex to keep pace with the delivery cadence, and work with the lead engineer to re-balance rather than let a quality backlog build up silently.
- Beyond the agents module, contribute to and help scale test automation practice across the wider --- platform - frontend E2E (Playwright) and backend (xUnit) - as the team's senior dedicated test engineer.
Required Experience & Skills
- 3-5+ years' experience in a software engineering or test automation role.
- Experience designing test strategy from scratch, not just executing an existing suite.
- Experience with C#/.NET testing (xUnit or similar) and/or JavaScript E2E testing frameworks (Playwright or similar).
- Comfortable reading and reasoning about LLM/agent output well enough to write meaningful test assertions against it.
- Ability to work independently across two engineers' agents in parallel, without becoming a bottleneck on their delivery cadence.
- Experience working in a sprint-based development team.
Desirable / Nice-to-Have Skills Experience with some of the following is beneficial but not required:
- Hands-on experience evaluating GenAI/LLM applications (golden sets, LLM-as-judge, eval frameworks such as Promptfoo, DeepEval, Ragas, or similar).
- Experience with Google Vertex AI, Gemini, or a comparable LLM platform.
- Exposure to MCPs, or tool-use integrations for LLMs.
- Familiarity with ad platform APIs (Google Ads, Microsoft Advertising, Meta, TikTok, DV360).
- Experience with prompt management tooling (we use PostHog).
- API testing experience (e.g. via Playwright's API testing capabilities).
Tech Setting The --- product is built using:
- Backend : C# (.NET 10), Microsoft Agent Framework, Google Vertex AI (Gemini)
- Frontend : ReactJS
- Data : MySQL, Firestore, Google Cloud Storage
- Infrastructure : Google Cloud Platform (Cloud Run, Cloud Functions, Cloud Scheduler, Docker)
- Prompt Management : PostHog
- Testing : xUnit (C#), Playwright (frontend E2E) - general repo conventions. The eval/golden-set framework for agent output does not exist yet and is what this role builds and owns
📌 AI QA Test Engineer (Bengaluru)
🏢 Dentsu Global Services
📍 Bengaluru