Agentic AI Research Engineer: Benchmarking & System Validation (Bengaluru)

Agentic AI Research Engineer: Benchmarking & System Validation (Bengaluru)

01 Oct
|
CraftifAI
|
Bengaluru

01 Oct

CraftifAI

Bengaluru

Company: CraftifAI

Product: Orbit

Education: Research-focused master’s degree required; PhD preferred.

About CraftifAICraftifAI is building Orbit, an agentic engineering platform that automates the development of embedded software, edge AI applications, and FPGA-based systems. Through PipeGen, FirmGen, and AgentIQ, Orbit connects requirements and system architecture with code generation, testing, debugging, and hardware validation.

Our focus goes beyond generating code: we aim to turn engineering intent into working, validated systems.

About the RoleWe are looking for a research-driven, hands-on Agentic AI Research Engineer to lead product benchmarking, system-level benchmark development, and validation planning and execution for CraftifAI Orbit.

You will define how we measure the quality, reliability, efficiency, and completeness of agentic engineering workflows and build the infrastructure to evaluate them reproducibly. Your work will answer a central question:

Can an AI-driven engineering workflow consistently turn a real-world requirement into a system that works correctly on its intended hardware, with measurable cost, effort, and reliability?

This is not a conventional QA or model-accuracy evaluation role. It combines applied research, experimental design, software engineering, and real-world system validation. You will own the evaluation approach, implement it, execute experiments, and translate findings into product improvements.

Key Responsibilities1.

Lead Product

Benchmarking for Orbit

- Define the benchmarking strategy across Orbit’s capabilities, including requirements interpretation, architecture generation, code and test generation, tool execution, debugging, and end-to-end task completion.
- Design and run controlled comparisons against relevant coding agents, alternative workflows, and engineer-led baselines. Establish transparent conditions for problem statements, hardware, toolchains, available tools, execution budgets, and human assistance.
- Measure meaningful engineering outcomes: functional success, requirement satisfaction, time to a validated solution, human intervention, token and compute cost, failure-recovery capability, and post-generation engineering rework. Analyze repeated runs, uncertainty, and failure patterns rather than relying on isolated demonstrations or a single aggregate score.

2.

Build

System-Level Benchmarks
- Create representative benchmark suites for embedded firmware, edge AI pipelines, and FPGA workflows in collaboration with domain experts.



Evaluate complete engineering tasks—from requirements and architecture to implementation, integration, and execution—not just isolated code snippets.
- Develop rigorous task definitions and evaluation methods, including reference implementations or independently verified expected behavior, acceptance criteria, scoring rubrics, difficulty levels, and requirement-to-test mappings. Include ambiguous requirements, integration challenges, hardware constraints, and failure conditions.
- Build reusable evaluation infrastructure for automated execution, toolchain integration, simulation, hardware interaction, result collection, and reporting. Maintain versioned datasets, held-out tasks, reproducible configurations, and benchmark histories that help distinguish genuine capability improvements from benchmark-specific tuning.

3. Plan and Execute Validation
- Own validation plans and acceptance criteria for agent-generated artifacts and complete systems. Cover unit, integration, system, simulation, and hardware-in-the-loop testing, as applicable, with traceability from requirements to test results.
- Execute functional, performance, and robustness validation on representative workloads and target hardware. Assess relevant measures such as application accuracy, end-to-end latency, throughput, memory use, power consumption, interface behavior, and recovery from faults or degraded operating conditions.
- Turn validation into a continuous product feedback loop. Automate regression testing and release checks, investigate failures with engineering teams, and verify improvements. Produce evidence-backed reports that clearly distinguish passed, failed, and untested behavior. Use independent checks so that generated code and generated tests do not simply reinforce the same incorrect assumptions.

Required Qualifications
- A master’s degree with a substantial research component in Computer Science, Artificial Intelligence, Electrical/Electronics Engineering, Robotics, or a related discipline. A PhD is preferred.
- Research or applied engineering experience in agentic AI, LLM evaluation, program synthesis, automated software engineering, or intelligent-system validation,



with evidence of designing and executing rigorous experiments.
- Strong Python and software engineering skills, including building automated experiment pipelines, evaluation tools, test harnesses, and reproducible development environments.
- Strong experimental and analytical ability: formulating hypotheses, selecting baselines, designing ablation studies, analyzing variability, and drawing conclusions supported by evidence.
- Ability to own an open-ended technical problem end to end, translate broad requirements into measurable criteria, investigate failures systematically, and communicate findings through transparent technical documentation.

Preferred ExperienceExperience in one or more of the following areas will be valuable; expertise across every hardware domain is not expected.
- Embedded and hardware-aware engineering: C/C++, microcontrollers, RTOS/Linux, edge AI deployment, FPGA/RTL workflows, simulators, or hardware-in-the-loop testing.
- Advanced evaluation and verification: long-horizon agent evaluation, tool-use assessment, fault injection, property-based testing, formal methods, hardware profiling, or automated measurement equipment.
- Demonstrated research contributions: publications, a strong research thesis, open-source evaluation frameworks, benchmark datasets, or reproducible projects in AI, software engineering, robotics, or embedded systems.

What Success Looks LikeYou will establish a trusted evaluation framework that shows what Orbit can do, where it fails, how reliably it performs, and whether new releases genuinely improve outcomes. Success means reproducible product comparisons, reusable system-level benchmarks, automated validation with clear acceptance evidence, and research findings that directly influence agent architecture, tool integration, and product development.

Why Join CraftifAI?Work on a question at the intersection of AI research and engineering: how to evaluate agents that must deliver working systems, not just convincing answers or compilable code. This role offers ownership of the benchmarking and validation function, close collaboration with product and engineering teams, and the opportunity to shape how agentic engineering capability is measured.

ApplicationPlease share your CV and relevant research or engineering work, such as publications, a thesis, GitHub repositories, benchmark projects, or evaluation frameworks. Highlight your specific contribution, experimental approach, and measurable results.

📌 Agentic AI Research Engineer: Benchmarking & System Validation (Bengaluru)
🏢 CraftifAI
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai research engineer: benchmarking & system validation (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai research engineer: benchmarking & system validation (bengaluru) / bengaluru