What you'll need to have done before
We are looking for evidence you have built and operated this kind of system, not that you have read about it.
- MCP servers you have authored yourself. Tool schema design, error surfaces, versioning, and the practical experience of watching a model misuse a tool because the description was ambiguous.
- Multi-agent orchestration in production. LangGraph, Pydantic AI, or a framework you built because the existing ones did not fit. We care that you can explain the state, retry and cancellation model, not which library you picked.
- Tool-calling design at scale. What changes when a model has access to a hundred-plus tools rather than a handful. Selection strategies, tool grouping, and how you keep the call budget sane.
- Prompt and context engineering as engineering. Versioned prompts, measured changes, retrieval that is tuned rather than hoped over. If your answer to a quality regression is to add another sentence to the system prompt and redeploy, this is not the right role.
- Structured-output discipline. Schema-constrained generation, validation, and repair loops. Comfort with the fact that a model will eventually emit something invalid and the system has to keep working.
- Robust backend engineering in a language of your choice. Python or TypeScript/Node are both fine, as is anything else you can argue for. Async concurrency, typing and testing matter more than the runtime.
Signals that would move you to the top of the list
- Text-to-SQL over a large, messy, real schema. Vanna, DDL-based retrieval, or your own approach.
- Enforcing authorisation inside generated queries, where the model must not be able to route around row-level access rules.
- Any exposure to SFT, LoRA or distillation. You will not own the training runs, but you will be a major producer of their input data, and knowing what makes a trajectory useful to train on matters.
- Eval design for non-deterministic systems. Deciding what "correct" means for a generated query or a patched procedure, and building a suite that stays honest as the model changes.
- Prompt injection and agent security thinking, particularly where an agent holds write access to production.
- T-SQL, stored procedures, or work inside a mature enterprise platform where the domain logic lives in the database rather than the application code.
- Channel integrations: WhatsApp Business API, Teams bot framework, Slack apps.
What this role is not
It is not prompt engineering against a hosted API. It is not building a RAG chatbot over a document set. It is not research, and you will not be training foundation models.
It is systems engineering in a domain where one of your dependencies is stochastic, the blast radius includes a production database, and every write is reviewed by a human before it lands. If that framing is appealing rather than tedious, we should talk.
📌 LLM Systems Engineer (Delhi)
🏢 Debound
📍 Delhi