31 Jul
|
LH2 AI Labs
|
Bengaluru
31 Jul
LH2 AI Labs
Bengaluru
About LH2 AI Labs
Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.
We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.
For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on — there is limited new signal left there. That is where we come in.
Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today. The Role
We are looking for an AI Research Engineer — AI Evaluation, Benchmarking & LLM Post-Training to help us build evaluation and training-data systems for frontier AI models and coding agents.
In this role, you will work at the intersection of software engineering, AI evaluation, benchmark design, and post-training. You will create realistic coding tasks from real repositories, build reproducible evaluation environments, design tests and verifiers, analyse model performance, and improve the quality of data used to train and evaluate AI systems.
This is a strong fit for a research-oriented software engineer with practical experience in coding benchmarks such as SWE-bench, SWE-Bench Pro, DeepSWE, Terminal-Bench, or equivalent repository-based evaluation frameworks. What You’ll Do Design realistic coding tasks covering debugging, feature development, refactoring, code review, repository navigation, and test-based problem solving. Source and structure tasks from repositories, issues, pull requests, commits, diffs, test suites, and customer-provided codebases.
Build reproducible evaluation environments using repositories, Docker containers, dependencies, test harnesses, sandboxes, and code-execution workflows. Define task instructions, acceptance criteria, reference solutions, grading rubrics, tests, hidden tests, and programmatic verifiers.
Run and adapt coding benchmarks such as SWE-bench, SWE-Bench Pro, DeepSWE, Terminal-Bench, or similar internal benchmarks. Evaluate outputs and trajectories generated by LLMs, humans, and coding agents, and determine whether tasks have been correctly solved. Identify ambiguous, broken, trivial, contaminated, non-reproducible,
or incorrectly verified tasks before they reach customers or training pipelines.
Analyse model and agent failure modes across repository understanding, planning, implementation, tool usage, debugging, and test execution. Build datasets for supervised fine-tuning, preference optimisation, RLHF, RLVR, and other LLM post-training workflows. Convert successful and unsuccessful agent trajectories into demonstrations, corrections, preference pairs, critiques, and outcome-labelled training examples.
Build and improve pipelines for collecting, validating, deduplicating, analysing, and packaging evaluation and post-training datasets. Work with customers and internal teams to convert high-level evaluation or training goals into clear benchmark workflows, technical environments, and quality standards. What We’re Looking For Graduate from a Tier 1 engineering institution such as IIT, BITS, NIT, IIIT, or equivalent. 3–7+ years of professional experience in software engineering, AI research engineering, machine-learning engineering, or developer tooling.
Strong programming ability in Python and at least one additional language such as JavaScript/ TypeScript, Java, Go, C++, or Rust. Robust ability to read, understand, modify, and debug unfamiliar production codebases.
Experience with Git, GitHub workflows, pull requests, issues, commits, diffs, branches, and code review.
Experience working with unit tests, integration tests, CI workflows, failing builds, dependency issues, and debugging logs. Ability to create and debug reproducible environments using Docker, Linux, shell scripts, package managers, and cloud or local execution setups. Hands-on experience building, running, adapting, or analysing coding-model or coding-agent benchmarks.
Practical knowledge of benchmarks such as SWE-bench, SWE-bench Verified, SWE-Bench Pro, DeepSWE, Terminal-Bench, or equivalent internal frameworks.
Experience creating evaluation tasks from repositories, issues, pull requests, commits, or real engineering requirements.
Experience designing test harnesses, hidden tests, reference solutions, grading rubrics, golden datasets,
or programmatic verifiers. Understanding of benchmark contamination, solution leakage, test overfitting, flaky environments, and reproducibility challenges. Ability to distinguish between model failures, agent failures, task-design problems, verifier problems, and environment failures.
Strong engineering judgment around correctness, reproducibility, code quality, task difficulty, and evaluation reliability. Strong written and verbal communication skills for documenting task design, methodology, results, and failure analysis. Nice to Have Experience building benchmarks or evaluation datasets for private repositories.
Experience with coding agents or developer tools such as SWE-agent, OpenHands, Claude Code, Cursor, Codex-style agents, or Devin-style systems.
Experience building datasets for RLHF, DPO, RLVR, verifier-based reinforcement learning, or other post-training workflows. Familiarity with additional benchmarks such as LiveCodeBench, HumanEval+, MBPP+, RepoBench, or Multi-SWE-bench.
Experience with secure code execution, sandboxing, automated grading, test generation, or distributed evaluation infrastructure. Prior experience working with AI labs, evaluation teams, data companies, or applied AI organisations. Success in This Role Looks Like Coding tasks are realistic, reproducible, well-scoped, challenging, and objectively verifiable.
Public and private repositories can be converted into reliable benchmark environments. Evaluation results accurately reflect model capabilities rather than problems in the task, verifier, or infrastructure.
Broken, ambiguous, contaminated, or low-quality tasks are identified early. Model failures are converted into useful benchmark improvements and high-quality post-training examples. Customers receive clean datasets, clear methodologies, meaningful analysis, and trustworthy evaluation results.
Why This Role
Matters
Training and evaluating AI coding systems requires more than collecting code or comparing generated patches with existing Git diffs.
Reliable benchmarks require realistic engineering tasks, reproducible repositories, strong tests, programmatic verifiers, secure execution environments, and careful human judgment.
This role will help LH2 AI Labs build the evaluation and post-training infrastructure required to improve the next generation of frontier AI models and autonomous software-engineering agents.
📌 AI Research Engineer - AI Evaluation, Benchmarking & LLM Post Training (Bengaluru)
🏢 LH2 AI Labs
📍 Bengaluru