- Develop and optimize AI inference kernels using C/C++.
- Optimize kernels for different data types including INT8, INT16, BF16, and FP32.
- Perform SIMD/vectorization, Intrinsics, Memory optimization, and Performance tuning.
- Profile and analyze application performance to identify optimization opportunities.
- Debug and resolve functional and performance issues.
- Integrate optimized kernels into the AI inference pipeline.
- Work closely with hardware and software teams to improve overall inference
Performance.
Requirements
- Valuable knowledge of C/C++ programming.
- Basic understanding of AI/Deep Learning models and inference.
- Experience with DSPs, hardware accelerators, or similar compute platforms.
- Knowledge of SIMD/vectorization, Intrinsics, and low-level software optimization.
- Good profiling, debugging,
and performance analysis skills.
- Understanding of computer architecture and memory optimization is an advantage.
Preferred Experience
- Experience in AI/ML inference optimization.
- Experience with C7x DSP, HWA-MMA, or similar hardware accelerators.
- Knowledge of different numerical data types such as INT8, INT16, BF16, FP32.
- Familiarity with performance profiling and benchmarking tools
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 AI/DSP Kernel Optimization Engineer (Chennai)
🏢 MulticoreWare
📍 Chennai
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.