13 Aug
|
NAVA
|
Bengaluru
Role & Responsibilities
- Design, write, and optimize high-performance CUDA kernels targeting NVIDIA Ampere, Hopper, and upcoming architectures.
- Profile and debug GPU memory bandwidth, occupancy, and instruction-level bottlenecks using Nsight Compute, Visual Profiler, and custom instrumentation.
- Collaborate with ML and HPC teams to translate algorithmic requirements into effective, scalable, and portable GPU code.
- Implement fused operations, shared memory optimizations, and warp-level primitives to maximize throughput under latency and memory constraints.
- Contribute to kernel library development and maintain reusable, well-documented, and performance-tested modules.
- Stay ahead of evolving CUDA toolchains, compiler flags,
and architecture-specific features to continuously extract performance wins.
Skills & Qualifications
Must-Have
- CUDA C++
- Nsight Compute
- NVIDIA GPU Architecture (Ampere/Hopper)
- GPU Memory Hierarchy Optimization
- Warp-Level Primitives
- Profiling & Performance Tuning
- nvcc Compiler
- PTX / SASS Assembly (Basic Understanding)
Preferred
- Experience with Tensor Cores and FP8/FP16 kernels
- familiarity with cuDNN, cuBLAS, or CUTLASS
- Contributions to open-source GPU compute projects
Skills: cuda,graphs,architecture,layout,design,contribute,fusion,kernel,c,computing
📌 CUDA Kernel Engineer (Bengaluru)
🏢 NAVA
📍 Bengaluru