You will design and build AI agents that do real work for our clients: answering support questions from their own documents, qualifying leads, screening candidates and handing over to a person at the right moment. This is production engineering, not research. Every agent you build ships to real users and has to keep working.
What you will do
- Build agents that answer from client data using retrieval, tools and clear guardrails.
- Connect agents to the systems clients already use, such as help desks, CRMs and WhatsApp.
- Create test sets from real conversations, and measure quality before and after every change.
- Monitor live agents, review conversations and improve answers over time.
- Explain trade-offs to clients in plain language.
You might be a valuable fit if
- You have shipped LLM-based features to production and know why some of them failed.
- You are comfortable in TypeScript or Python, and with APIs, queues and databases.
- You care about evaluation, not just demos.
- You can work in weekly increments and enjoy showing your work.
Nice to have
- Experience with retrieval, embeddings and vector or hybrid search.
- Experience with Zendesk, Intercom or the WhatsApp Business Platform.
How we work
- A remote-first team, working with clients in the US, UK and India.
- We ship working software every week and show it to the client.
- Clients own the source code and infrastructure we build, so we earn our next project by doing this one well.