You will design and build AI agents that do real work for our clients: answering support questions from their own documents, qualifying leads, screening candidates and handing over to a person at the right moment. This is production engineering, not research. Every agent you build ships to real users and has to keep working.
What you will do
Build agents that answer from client data using retrieval, tools and transparent guardrails.
Connect agents to the systems clients already use, such as help desks, CRMs and WhatsApp.
Create test sets from real conversations, and measure quality before and after every change.
Monitor live agents, review conversations and improve answers over time.
Explain trade-offs to clients in plain language.
You might be a valuable fit if
You have shipped LLM-based features to production and know why some of them failed.
You are comfortable in TypeScript or Python, and with APIs, queues and databases.
You care about evaluation, not just demos.
You can work in weekly increments and enjoy showing your work.
Nice to have
Experience with retrieval, embeddings and vector or hybrid search.
Experience with Zendesk, Intercom or the WhatsApp Business Platform.
How we work
A remote-first team, working with clients in the US, UK and India.
We ship working software every week and show it to the client.
Clients own the source code and infrastructure we build, so we earn our next project by doing this one well.