12 Sep
|
Lyzr AI
|
Bengaluru
Job Description
Model Trainer (Post-Training and Evals)
nThe job
nTake an agentic task that currently runs on a frontier model, fine-tune an open-weight model on it, and prove with evals that the small model does the task better and cheaper. Then do it again on the next task, and the next.
nYou will work on real enterprise workloads. Some of it runs on customer GPUs, some on ours. You own the model from data through to the eval report that convinces a customer to switch traffic over.
nWhat you will do
n
- n
- Build training sets from real agent traffic and traces, not scraped benchmarksn
- Run LoRA, QLoRA, DoRA, full fine-tunes, CPT, DPO and GRPO depending on what the task needsn
- Design task-specific evals for agentic behavior: tool calling accuracy, multi-step reliability, output format adherence, groundedness. Benchmarks are not the bar, the customer's task isn
- Get the model serving on vLLM and measure real latency and token cost against the frontier baselinen
- Work inside ShadowLM, our open-source fine-tuning SDK, and improve it as you gon
nWhat we are looking for
n
- n
- You have post-trained open-weight models yourself and can talk about what failed, not just what workedn
- Comfortable across llama, qwen, mistral, gemma, phi, deepseek and whatever ships next monthn
- Solid on evals. You know that most fine-tuning projects die because nobody built an honest measurement firstn
- You can read a paper on Monday and have it running on Wednesdayn
- Practical about hardware. You know what fits on one GPU and what does notn
nNice to have
n
- n
- Reward modeling or RL for agentsn
- Quantization and distillationn
- Experience getting a model past an enterprise security reviewn
📌 Model Trainer (Bengaluru)
🏢 Lyzr AI
📍 Bengaluru