28 Sep
|
MulticoreWare
|
India
28 Sep
MulticoreWare
India
Job Information
- Department Name Frameworks & Cloud
- Job Type Full time
- Date Opened 30/08/2026
- Industry Software Development
- Minimum Experience In Years 5
- Maximum Experience In Years 10
- City Ramapuram
- Province Tamil Nadu
- Country India
- Postal Code 600089
About Us
MulticoreWare is a global software solutions & products company with its HQ in San Jose, CA, USA. With worldwide offices, it serves its clients and partners in North America, EMEA and APAC regions. Started by a group of researchers, MulticoreWare has grown to serve its clients and partners on HPC & Cloud computing, GPUs, Multicore & Multithread CPUS, DSPs, FPGAs and a variety of AI hardware accelerators.
MulticoreWare was founded by a team of researchers that wanted a better way to program for heterogeneous architectures. With the advent of GPUs and the increasing prevalence of multi-core, multi-architecture platforms, our clients were struggling with the difficulties of using these platforms efficiently.
We started as a boot-strapped services company and have since expanded our portfolio to span products and services related to compilers, machine learning, video codecs, image processing and augmented/virtual reality. Our hardware expertise has also expanded with our team; we now employ experts on HPC and Cloud Computing, GPUs, DSPs, FPGAs, and mobile and embedded platforms. We specialize in accelerating software and algorithms, so if your code targets a multi-core, heterogeneous platform, we can help.
Job Description
We are looking for a Senior Software Engineer to develop and optimize deep learning models, including CNNs, LLMs, and MoE, for efficient inference across CPU, GPU, hardware accelerators, and edge devices. The role focuses on quantization, model compression, high-performance kernel implementation, transformer optimization, and production deployment.
Responsibilities:
-
Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, hardware accelerators, and edge devices.
-
Design and implement quantization algorithms (PTQ, QAT, GPTQ, AWQ) from scratch.
-
Apply model compression techniques such as pruning, decomposition, and distillation.
-
Implement and optimize quantized kernels (INT8, INT4, FP8) using C++ for high performance.
-
Translate research papers into production-ready implementations.
-
Optimize latency, throughput, and memory usage for real-world deployment.
-
Work on transformer optimization including KV-cache, PEFT (LoRA/QLoRA), and MoE models.
-
Profile, benchmark, and debug model performance across different hardware platforms.
-
Collaborate with ML, compiler, and hardware teams to deliver optimized solutions.
Requirements
Education:
BE/BTech/MS/MTech in Computer Science or a related field.
Technical Skills (Must haves):
- 4+ years of relevant experience.
-
Strong programming skills in Python and C++.
-
Proven experience in quantization algorithms (PTQ, QAT, GPTQ, AWQ).
-
Hands-on experience in pruning, model compression, and inference optimization.
-
Experience implementing quantization or optimization techniques from scratch.
-
Robust understanding of CNNs, Transformers, and LLM architectures.
-
Experience with PyTorch / ONNX and model deployment pipelines.
-
Strong problem-solving and performance optimization skills.
Need to have (Can be bridged):
No additional bridged skills were specified.
Good to have (Not essential):
-
Experience with MoE architectures and PEFT techniques (LoRA, QLoRA).
-
Knowledge of TensorRT, ONNX Runtime, TVM, and MLIR.
-
Familiarity with hardware-aware optimization across GPU, NPU, and edge devices.
-
Experience in research paper implementation or open-source contributions.
Preferred Qualifications (Optional):
No additional preferred qualifications were specified.
📌 SSE - Optimization Engineer (India)
🏢 MulticoreWare
📍 India