06 Aug
|
Advanced Micro Devices (AMD)
|
Hyderabad
06 Aug
Advanced Micro Devices (AMD)
Hyderabad
THE TEAM:
Join AMD's high-impact team at the heart of innovation in AI, ML, and high-performance computing (HPC). We're a team-oriented group of software architects and GPU engineers focused on pushing the boundaries of AI model performance across distributed, GPU-accelerated platforms. Our work drives the next generation of AMD's AI software stack, enabling large-scale machine learning training and inference workloads in data centers and enterprise environment.
THE ROLE:
AMD is hiring a Senior Software Developer for its AI, ML, and high-performance computing team to lead GPU kernel optimization and distributed software for large-scale AI workloads. In this technical leadership role, you'll architect and implement optimized compute kernels, design multi-GPU/multi-node scaling strategies, profile systems to maximize hardware utilization, build benchmarking infrastructure, and guide agile teams across the product lifecycle. The ideal candidate is a deep systems thinker fluent in GPU architecture, parallel computing, and AI model execution, comfortable both writing performance-critical code and driving software architecture decisions. Required expertise includes GPU kernel optimization in C++ (17/20), hands-on CUDA and low-level GPU programming, distributed AI computing (multi-GPU, NCCL, MPI), and familiarity with frameworks such as PyTorch, vLLM, Cutlass, and Kokkos, plus strong performance tuning with profiling tools (Nsight, VTune, Perf) and Python automation. You should also bring proven software leadership experience defining roadmaps with stakeholders, interfacing with executives, and shipping production software through open-source upstreaming or commercial rollouts.
THE PERSON:
We're looking for a highly skilled, deep systems thinker who thrives in complex problem domains involving parallel computing, GPU architecture, and AI model execution.
You are confident leading software architecture decisions and know how to translate business goals into robust, optimized software solutions. You're just as comfortable writing performance-critical code as you are guiding agile development teams across product lifecycles. Ideal candidates have a strong balance of low-level programming, distributed systems knowledge, and leadership experience paired with a passion for AI performance at scale.
KEY RESPONSIBILITIES:
- GPU Kernel Optimization: Develop and optimize GPU kernels to accelerate inference and training of large machine learning models while ensuring numerical accuracy and runtime efficiency.
- Multi-GPU and Multi-Node Scaling: Architect and implement strategies for distributed training/inference across multi-GPU/multi-node environments using model/data parallelism techniques.
- Performance Profiling: Identify bottlenecks and performance limitations using profiling tools; propose and implement optimizations to improve hardware utilization.
- Parallel Computing: Design and implement multi-threaded and synchronized compute techniques for scalable execution on modern GPU architectures.
- Benchmarking & Testing: Build robust benchmarking and validation infrastructure to assess performance, reliability, and scalability of deployed software.
- Documentation & Best Practices: Produce technical documentation and share architectural patterns, code optimization tips, and reusable components.
PREFERRED EXPERIENCE:
- GPU kernel development (HIP, CUDA C/C++, PTX, GPU Assembly)
- GPU kernel optimization down to assembly level
- GPU hardware architectures (AMD, nVidia, Intel)
- ML/HPC related parallel algorithm design, e.g., GEMMs, element-wise, attention, reductions
- ROCm/CUDA Software Stacks (Runtimes, Compilers, Libraries)
- Scripting knowledge (ex. Python)
- Advanced C++ software development (including meta-programming, C++20 features)
- Software development, analyzing, and debugging of complex algorithms.
- Reading, understanding, and changing advanced C++
- Reading, understanding, and changing complex assembly
- Advanced knowledge of software development processes.
- Excellent working knowledge of GPU/CPU architectures.
- Distributed computing and multi-GPU environments.
- Advanced performance profiling and optimization tools.
- C++ Performance optimization
- Optimizing GPU kernels in C++20.
- Strong experience in low-level GPU kernel optimization.
- Optimization of GPU assembly
- Practical usage of the LLVM compiler flow and tools
- Proficiency in HIP / CUDA and GPU programming.
- GPU performance bottleneck analysis
- GPU Power Optimization analysis
- Understanding Neural Network models data flow and operations
- Working with complex frameworks/libraries such as (PyTorch, vLLM, CUTLASS, Kokkos, etc.)
- Working on complex algorithms at operator level (variances of attention algorithms,MoE, quantization/scaling etc.)
- OS Kernel Debug and Kernel optimization
- GPU API programming
- CPU/ASIC software development
ACADEMIC CREDENTIALS
- Bachelor's or Master's degree in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field.
- Advanced degrees or published work in HPC, GPU computing, or AI systems is a plus.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Lead HPC Software Optimization Engineer - C++ (Hyderabad)
🏢 Advanced Micro Devices (AMD)
📍 Hyderabad