19 Sep
|
Gradientflo Labs
|
Hyderabad
19 Sep
Gradientflo Labs
Hyderabad
Role Overview
You’ll be the safety net and enabler of speed at Vibecoderz. In a system where AI agents generate artifacts, mini-apps, and real-time interactions, quality isn’t just about correctness — it’s about trust.
As the founding QA Automation Engineer, you’ll design frameworks that validate not just code but also AI outputs. From traditional E2E flows to specialized agent evaluation pipelines, you’ll ensure that everything shipping to production meets our gold standard.
This role is about building autonomous, self-healing QA systems that scale with our multi-agent platform. You’ll work hand-in-hand with PM, FE, BE, AI, and DevOps to embed quality into every phase of the pipeline using Linear (execution), Notion (test cases/playbooks), and GitHub (CI/CD integration) as your tools of record.
Key Responsibilities
1. Automation Framework Ownership
- Design and implement end-to-end automated test frameworks for FE, BE, and AI agents.
- Cover unit, integration, regression, and smoke tests.
- Frontend Testing
- Build E2E suites for Next.js UI using Playwright or Cypress.
- Simulate real-user flows: onboarding, AI tutor interaction, artifact generation.
- Backend Testing
- Automate FastAPI endpoint validation.
- Ensure <100ms response SLA is consistently met.
- AI Agent Evals
- Develop golden sets for agent responses (TutorAgent, Mini-App Generator, Quiz Generator).
- Automate detection of hallucinations, irrelevant answers, or unsafe outputs.
- Continuous Testing Integration
- Integrate all test suites into GitHub Actions CI/CD pipelines.
- Block merges unless all critical tests pass.
- Realtime Testing
- Validate WebSocket-based real-time interactions (voice transcription, live artifacts).
- Create stress tests for multi-user editing and presence.
- Security & Reliability Testing
- Build automated checks for prompt injection and unsafe code generation.
- Implement sandboxed test environments for mini-app execution.
- Performance & Load Testing
- Benchmark system under 100K concurrent learners.
- Identify latency bottlenecks across Pub/Sub and Cloud Run services.
- Defect Tracking & Reporting
- Use Linear for bug triage and prioritization.
- Create dashboards of defect density, test coverage, and agent accuracy.
- Culture of Quality
- Define QA best practices, playbooks, and documentation in Notion.
- Mentor engineers to write better testable code.
Success Metrics
90 Days (Probation):
- Automated framework live for FE (Next.js) and BE (FastAPI).
- AI TutorAgent “Text → Course” flow covered by golden set evals.
- CI/CD pipeline enforces automated test runs before deploy.
12 Months:
- 85% test automation coverage across platform.
- <1% escaped defects in production.
- AI agent eval accuracy >90% on golden set outputs.
- System proven under 100K concurrent users with <1% error rate.
Must-Haves
- 10+ years QA automation in SaaS/product engineering.
- Solid expertise with Playwright/Cypress, Pytest, Postman/Newman.
- Proven experience integrating test suites into CI/CD pipelines.
- Familiarity with real-time systems (WebSockets, CRDT/OT).
- Deep understanding of QA methodologies: smoke, regression, load, E2E.
Nice-to-Haves
- Experience with LLM/AI output evaluation (evals).
- Prior startup or founding engineer experience.
- Knowledge of security testing (XSS, prompt injection defense).
- Contributions to open-source QA frameworks.
Tech Stack Visibility
- Frontend: Playwright, Cypress
- Backend: Pytest, Postman/Newman
- AI Evals: Custom golden set framework, LangSmith/Langfuse for prompt monitoring
- CI/CD: GitHub Actions integration
- Observability: OpenTelemetry hooks into test reports
- Infra: Cloud Run staging environments for test runs
Assessment
Objective: Validate ability to design automated test frameworks that cover FE, BE, and AI outputs.
Challenge (Candidate PoC):
1. Frontend (FE) E2E Test
- Write Playwright/Cypress tests for:
- User onboarding (sign-up with Google).
- Chatting with TutorAgent mock API.
- Rendering a Quiz block with answer validation.
- Backend (BE) API Test
- Write Pytest suite for FastAPI service:
- /generate-course endpoint (input: “Teach React Hooks”).
- Assert valid JSON structure for output.
- AI Agent Eval
- Create a golden set with 5 input prompts and expected TutorAgent responses.
- Automate comparison with tolerance for natural language variation.
- CI/CD Integration
- Configure GitHub Actions to run FE + BE + AI eval tests on each PR.
- Block merge if any critical tests fail.
Deliverables:
- GitHub repo with FE, BE, and AI test suites.
- GitHub Actions workflow file with integrated tests.
- Golden set JSON for TutorAgent eval.
- README documenting design choices and edge cases.
Evaluation Criteria:
- Automation Framework Quality (30%)
- AI Evals & Innovation (25%)
- CI/CD Integration (20%)
- Test Coverage & Reliability (15%)
- Documentation & Clarity (10%)
📌 QA Automation Engineer (Founding Team) (Hyderabad)
🏢 Gradientflo Labs
📍 Hyderabad