02 Sep
|
MulticoreWare
|
Chennai
02 Sep
MulticoreWare
Chennai
Title - Senior Software Engineer
Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, and hardware accelerators / edge devices.
- Design and implement Quantization algorithms (PTQ, QAT, GPTQ, AWQ) from scratch.
- Apply model compression techniques such as pruning, decomposition, and distillation.
- Implement and optimize quantized kernels (INT8, INT4, FP8) using C++ for high performance.
- Translate research papers into production-ready implementations.
- Optimize latency, throughput, and memory usage for real-world deployment.
- Work on transformer optimization including KV-cache, PEFT (LoRA/QLoRA), and MoE models.
- Profile, benchmark, and debug model performance across different hardware platforms.
- Collaborate with ML, compiler, and hardware teams to deliver optimized solutions.
Must-Have
- BE/BTech/MS/MTech in Computer Science or related field with 4+ years of experience.
- Strong programming skills in Python and C++.
- Proven experience in Quantization algorithms (PTQ, QAT, GPTQ, AWQ).
- Hands-on experience in pruning, model compression, and inference optimization.
- Experience implementing quantization or optimization techniques from scratch.
- Solid understanding of CNNs, Transformers, and LLM architectures.
- Experience with PyTorch / ONNX and model deployment pipelines.
- Strong problem-solving and performance optimization skills.
Nice-to-Have
- Experience with MoE architectures, and PEFT techniques (LoRA, QLoRA).• Knowledge of TensorRT, ONNX Runtime, TVM, MLIR.
- Familiarity with hardware-aware optimization (GPU, NPU, Edge Devices).
- Experience in research paper implementation or open-source contributions.
📌 SSE - Optimization Engineer (Chennai)
🏢 MulticoreWare
📍 Chennai