You'll work directly with our CTO on multi-agent LLM systems: agents that parse real-world documents, reason over structured data, and produce validated outputs under a safety layer. Concrete work includes building and testing agent pipelines, prompt engineering with proper evaluation, retrieval (RAG) over domain data, and writing the guardrails that keep agent outputs correct and secure.
You should have
Hands-on experience building LLM agents not just calling an API, but chaining steps, tool use, and handling failure modes
Solid Python; comfort with FastAPI or similar for exposing agent workflows
Practical familiarity with at least one agent/LLM framework (LangChain, LlamaIndex, or equivalent) and RAG concepts
Enough understanding of evaluation to answer "how do you know your agent is right?" with something better than "it looked fine"
A project, repo,
or internship where you actually shipped something agent-based (link it)
Bonus
Experience with output validation, structured extraction, or explainability (e.g., SHAP)
Exposure to healthcare, clinical, or other high-stakes data domains
CI/CD, testing discipline, Git fluency
Who thrives here
We're a tiny team, so we care as much about how you work as what you know. We're looking for someone who takes feedback well, asks questions when stuck instead of hiding it, documents their work so others can build on it, and is genuinely team-oriented. Robust opinions are welcome; ego is not.
What you'll get
Direct mentorship from a senior founding team
Ownership of real features in a shipping product
A potential pre-placement offer for those who excel