- Experience shipping and operating an LLM or agent system used by real customers.
- Strong software engineering skills in TypeScript or Python, with ability to work across both.
- Experience building evaluations, datasets, experiments, or AI quality systems.
- Strong product judgment and ability to turn vague AI quality issues into measurable problems.
- Ability to work across data, evaluation methods, model selection, and fine-tuning.
- Strong ownership as a senior individual contributor.
- Bonus: Experience with multimodal AI, video, media, or creative software.
- Bonus: Experience with human labeling, model graders, or fine-tuning.
- Bonus: Strong understanding of experiment design and statistics.
Responsibilities
- Define quality standards for Cardboard’s agent.
- Build trusted evaluation datasets from real product usage.
- Build offline and online evaluations using automated checks, model graders, and human review.
- Analyze real agent runs and identify recurring failure patterns.
- Improve agent quality through better data, evaluation methods, model selection, and fine-tuning.
- Build regression checks and release gates for significant agent changes.
- Track AI quality alongside latency and cost.
- Partner with product and engineering teams to ship measurable improvements.
Job Details
Bengaluru, India
Interview Process
- Recruiter Screen
- Technical Interview
- ML & Evaluation Deep Dive
- Product & Engineering Interview
- Final Interview