29 Aug
|
Zenwork
|
Hyderabad
Dear Candiate
We are hiring for the below position.
Please go through the Job summary and acknowledge the same for further process.
Role: QA Automation AI Testing
Experience : 3+ Years
Location: Hyderabad
Notice Period: Immediate / 30 Days
Work Mode: Work From Office
Job Summary :-
What You Will Bring
- 3 years in QA with hands-on test automation: Python (Pytest) or JavaScript (Playwright or Cypress). This is the core of the role.
- API testing experience (Postman, REST) and comfort with Git and CI pipelines.
- A working understanding of LLM behavior: non-determinism, prompt sensitivity, and hallucination, with hands-on exposure to testing or evaluating an LLM feature, chatbot, or personal AI project.
- The ability to design test datasets and rubrics for outputs that have no single correct answer.
- Solid exploratory testing instincts and crisp bug reports with reproduction steps and evidence.
- SQL to verify data pipeline outputs and scores.
- Exposure to LLM evaluation tooling such as Promptfoo, LangSmith, or DeepEval is a plus, as is fintech, tax, or compliance domain experience.
The Role
AI systems do not behave like traditional software.
The same input can produce different outputs, quality is judged on rubrics rather than exact matches, and a prompt or model change can silently shift behaviour across the product.
We are hiring a QA Engineer who can test both worlds: solid automation across APIs and UI, plus structured evaluation of AI outputs for accuracy, consistency, safety, and drift.
You will be the quality gate for AI features that customers and internal teams rely on during the most time-sensitive weeks of the tax year.
You will build automated suites, design golden datasets and evaluation rubrics for LLM outputs, and stay hands-on with exploratory testing of conversational and agentic flows that automation alone cannot cover. This is a role for someone who does both automation and manual testing well, not one or the other.
What You Will Do
- Design, build, and maintain automated test suites across API, UI, and AI-output evaluation harnesses.
- Test conversational AI end to end: chat and calling flows, intent handling, resolution quality, human-handover triggers, and sentiment tagging accuracy.
- Validate agentic workflows: multi-step agent actions such as case creation, context assembly, and recommended actions, confirming the right steps happen in the right order with the right data.
- Build golden datasets and pass/fail rubrics for model outputs such as summaries, transcripts, scores, and classifications, including hallucination and accuracy checks against ground truth.
- Run prompt and model regression cycles: replay evaluation suites on every prompt or model change and report drift with evidence.
- Perform structured manual and exploratory testing of AI conversations and edge cases automation cannot reach.
- Conduct PII leakage checks in transcripts and outputs, and sample tax-content correctness in AI responses.
- Own bug triage, curate test datasets, and keep test documentation current.
📌 QA Automation AI Testing (Hyderabad)
🏢 Zenwork
📍 Hyderabad