Inference Server Engineer (Bengaluru)

Inference Server Engineer (Bengaluru)

19 Aug
|
Evollabs Tech
|
Bengaluru

19 Aug

Evollabs Tech

Bengaluru

Description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand LLM inference servers, model execution runtimes, MoE inference, and distributed AI systems. This role is focused on integrating our AI accelerator platform as a backend into contemporary LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar serving systems.

You will work on adding backend support for our hardware, enabling dense and MoE LLM inference, integrating custom operators and runtime paths,



and optimizing the execution of large-scale models across NPU and heterogeneous NPU-GPU systems. You will work at the intersection of LLM serving, accelerator runtime, model execution, distributed inference, and hardware-software co-design.

Responsibilities

Integrate our AI accelerator backend into LLM inference servers such as vLLM, SGLang, TensorRT-LLM, or similar frameworks.
Implement backend support for device registration, runtime execution, memory management, custom operators, and KV cache handling.
Enable dense and MoE LLM inference on our hardware, including attention execution, expert routing, expert execution, and distributed inference support.
Profile, debug, and optimize inference performance while building correctness, performance, and stress tests for NPU and heterogeneous NPU-GPU deployments.

Requirements

5+ years of relevant experience with M.S./Ph.D. degree in CS/CE or equivalent experience.
Strong C/C++ and Pyt

📌 Inference Server Engineer (Bengaluru)
🏢 Evollabs Tech
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: inference server engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: inference server engineer (bengaluru) / bengaluru