Greeting from TEKNIKOZ
Location- Remote
Experience- 10-12yrs or +
We are looking for a Compiler & Performance Engineer to develop the compiler stack for our custom AI accelerators and optimize contemporary Transformer/LLM workloads on our hardware.
Responsibilities
Develop compiler infrastructure, IRs, lowering passes, code generation, and hardware-specific optimizations.
Optimize Transformer/LLM workloads including GEMM, attention, normalization, KV-cache, and data movement.
Develop and optimize low-level kernels for our accelerator architecture.
Profile workloads and identify bottlenecks across compute, memory, and communication.
Improve latency, throughput, memory efficiency, and accelerator utilization.
Work closely with hardware, ML, runtime,
and systems teams on hardware-software co-design.
Enable recent ML models and operators on the accelerator.
Requirements
Strong C++ and Python programming skills.
Strong understanding of compiler fundamentals and computer architecture.
Experience with performance optimization, profiling, and low-level systems.
Familiarity with ML compilers such as LLVM/MLIR, TVM, XLA, or Triton is a plus.
Experience with Transformers/LLMs, GPU/NPU programming, or AI accelerators is highly desirable.
Solid analytical and debugging skills.
📌 Compiler & Performance Engineer : Ai Accelerators Pune
🏢 teknikoz
📍 Pune