06 Aug
|
Qpisemi
|
Bengaluru
AI20P Library Engineer, Machine Learning Acceleration
About the Role
We are looking for an experienced AI20P Library Engineer to bridge the gap between cutting-edge AI research and our hardware accelerators. In this role, you will design, develop, and optimize high-performance software libraries, kernels, and parallel compute runtimes for distributed AI20P environments. You will empower ML researchers and cloud developers to train and deploy frontier AI models at unprecedented scales.
Key Responsibilities
Kernel Development & Optimization: Design, implement, and tune high- performance custom operators and mathematical kernels specifically for AI20P architectures (assembly or intrinsic level).
Distributed Computing Runtimes: Build software abstractions and libraries that manage multi-host setups, sharding, and multi-dimensional parallelization (e.g., Megatron-LM style tensor/pipeline parallelism).
Communication Primitives: Design and optimize custom collective communication algorithms (AllReduce, AllGather) to minimize latency and maximize throughput over high-speed interconnects.
Performance Tuning: Profile distributed training workloads to identify and eliminate memory bottlenecks, network stalls, and suboptimal hardware utilization.
Cross-Functional Collaboration: Partner with compiler developers (XLA), hardware architects, and research scientists to co-design the future software- hardware ecosystem.
Qualifications
Education: B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a highly quantitative field.
Programming: Exceptional proficiency in C++ and Python.
ML Ecosystem: Hands-on experience with at least one major AI/ML framework (e.g., JAX, PyTorch).
Systems Knowledge: Solid understanding of concurrent computations, memory hierarchies (HBM, cache), and hardware accelerators.
Communication: Excellent cross-functional communication skills to translate complex research needs into production-ready software architecture.
Preferred Qualifications (Nice-to-Have)
Experience working with compiler construction or optimizing code generation for hardware.
Familiarity with cloud-based cluster managers (e.g., Kubernetes, SLURM) for executing large-scale distributed ML workloads.
Prior contributions to optimizing Large Language Models (LLMs) or multimodal architectures for low-precision inference and training.
📌 P Library Engineer, Machine Learning Acceleration (Bengaluru)
🏢 Qpisemi
📍 Bengaluru