? Responsibilities: Run evaluation sets, review AI outputs, curate ground truth, identify root causes, and provide actionable feedback to engineering teams.
⭐ Plus: Experience with Langfuse, Arize Phoenix, or similar AI evaluation/tracing tools
?️ Robust English communication and willingness to overlap with US working hours required.
? Interested candidates can apply using the link below: