02 Sep
|
Solumn AI
|
India
Urgent hiring alert! Immediate start! Fully Remote!
We’re looking to hire Technical Team Leads with solid hands-on experience in Agentic AI, LLM Engineering and evaluation data creation. The core requirement is experience setting up tool-calling environments where AI agents interact with APIs, applications and structured tools in realistic workflows. Experience with Python, FastAPI, MCP/A2A, function calling, Docker, sandboxing, JSON schemas, LangGraph, LangChain, CrewAI, AutoGen, RAG and cloud platforms would be highly relevant.
A major part of the role involves creating high-quality RLE and evaluation datasets for AI agents. This includes designing system and user prompts, defining expected states, forbidden actions and safety constraints, creating tool definitions and tool-call trajectories,
and building realistic episodes that test whether an agent can complete a task correctly and safely.
Experience with LLM evaluations, graders/verifiers, trajectory analysis, RLHF/SFT, AI safety, red teaming, prompt injection, SWE-Bench/SWE-Lancer-style benchmarks and human-in-the-loop workflows is a strong plus.
We are specifically looking for people who have led engineering teams while continuing to remain hands-on technically. You should be comfortable owning environment setup, dataset quality, code reviews, debugging, evaluation logic and team delivery. If you have built agentic evaluation environments, tool-use benchmarks, RLE datasets or similar systems, please reach out or share your profile.
📌 Team Lead - Agentic Tooling (India)
🏢 Solumn AI
📍 India