We are seeking expert-level GPU Software Engineers to support a high-visibility platform, focused on building software tooling on top of a custom compiler and SDK.
The role involves developing, optimizing, and porting GPU kernels and AI workloads to a specialized hardware platform.
Key Responsibilities
- Develop high-performance GPU kernels for specialized hardware platforms using PyTorch and Triton frameworks
- Build software solutions leveraging custom compiler toolchains and SDK capabilities
- Design and implement kernel-level optimizations to control and improve hardware execution behavior
- Port and optimize open-source AI/ML models for proprietary or custom SDK environments
- Adapt and run high-performance computing benchmarks and stress workloads, including:
- High Performance Linpack (HPL)
- BERT and similar benchmark-style workloads
- Develop stress testing and validation workloads aligned with hardware behavior and platform reliability goals
- Support validation and performance testing for both current and next-generation hardware platforms
- Collaborate closely with platform architects, compiler teams, and system engineers to enhance end-to-end system performance
Core Technical Skills (Must-Have)
Programming & Frameworks
- Strong proficiency in Python
- Expertise in C/C++ for systems-level programming
- Hands-on experience with PyTorch
- Experience with Triton (kernel development and Triton language)
GPU & Systems Expertise
- Mandatory experience in GPU kernel development
- Robust understanding of GPU architecture and performance optimization techniques
- Experience with compiler optimizations and runtime execution models
- Familiarity with custom SDKs and hardware abstraction layers
Performance & Workloads
- Hands-on experience in:
- GEMM (matrix multiplication) kernel development
- Porting and optimizing ML models on new/custom hardware platforms
- System-level performance tuning, benchmarking, and stress testing