Job Description
Role Overview
We are looking for a Software Lead (8+ years’ experience) to own the runtime and neural network (NN) layer of a next-generation AI accelerator platform. This role focuses on designing, optimizing, and implementing NN operators and developing current ops using CUDA/custom runtime APIs to deliver high-performance execution on custom AI hardware.
Key Responsibilities
Design and optimize NN operators for performance-critical workloads
Develop recent NN ops using CUDA/custom runtime APIs
Drive runtime-level optimizations across compute, memory, and scheduling
Own runtime ↔ NN layer interfaces and execution model
Implement and optimize operator fusion (e.g., matmul + bias + LayerNorm) for productive hardware utilization
Identify and resolve performance bottlenecks across the stack
Collaborate with compiler, PyTorch framework, and low-level SW teams
Impact
Own how efficiently AI workloads execute on the platform
Drive performance, scalability, and hardware utilization through optimized runtime and NN ops design
📌 Ai Software Lead – Pytorch & Cuda Runtime Next Gen Accelerator 10 Years Bengaluru (India)
🏢 Sandisk
📍 India