06 Aug
|
NAVA
|
Bengaluru
Role & Responsibilities
Design, write, and optimize high-performance CUDA kernels targeting NVIDIA Ampere, Hopper, and upcoming architectures.
Profile and debug GPU memory bandwidth, occupancy, and instruction-level bottlenecks using Nsight Compute, Visual Profiler, and custom instrumentation.
Collaborate with ML and HPC teams to translate algorithmic requirements into effective, scalable, and portable GPU code.
Implement fused operations, shared memory optimizations, and warp-level primitives to maximize throughput under latency and memory constraints.
Contribute to kernel library development and maintain reusable, well-documented, and performance-tested modules.
Stay ahead of evolving CUDA toolchains, compiler flags,
and architecture-specific features to continuously extract performance wins.
Skills & Qualifications
Must-Have
CUDA C++
Nsight Compute
NVIDIA GPU Architecture (Ampere/Hopper)
GPU Memory Hierarchy Optimization
Warp-Level Primitives
Profiling & Performance Tuning
nvcc Compiler
PTX / SASS Assembly (Basic Understanding)
Preferred
Experience with Tensor Cores and FP8/FP16 kernels
familiarity with cuDNN, cuBLAS, or CUTLASS
Contributions to open-source GPU compute projects
Skills: cuda,graphs,architecture,layout,design,contribute,fusion,kernel,c,computing
📌 CUDA Kernel Engineer (Bengaluru)
🏢 NAVA
📍 Bengaluru