You'll work directly with our CTO on multi-agent LLM systems: agents that parse real-world documents, reason over structured data, and produce validated outputs under a safety layer. Concrete work includes building and testing agent pipelines, prompt engineering with proper evaluation, retrieval (RAG) over domain data, and writing the guardrails that keep agent outputs correct and safe.
You should have
- Hands-on experience building LLM agents not just calling an API, but chaining steps, tool use, and handling failure modes
- Solid Python; comfort with FastAPI or similar for exposing agent workflows
- Practical familiarity with at least one agent/LLM framework (LangChain, LlamaIndex, or equivalent) and RAG concepts
- Enough understanding of evaluation to answer "how do you know your agent is right?" with something better than "it looked fine"
- A project, repo,
or internship where you actually shipped something agent-based (link it)
Bonus
- Experience with output validation, structured extraction, or explainability (e.g., SHAP)
- Exposure to healthcare, clinical, or other high-stakes data domains
- CI/CD, testing discipline, Git fluency
Who thrives here
We're a tiny team, so we care as much about how you work as what you know. We're looking for someone who takes feedback well, asks questions when stuck instead of hiding it, documents their work so others can build on it, and is genuinely collaborative. Robust opinions are welcome; ego is not.
What you'll get
- Direct mentorship from a senior founding team
- Ownership of real features in a shipping product
- A potential pre-placement offer for those who excel