- 8–10 years in software engineering, with 3+ years in on-device / edge AI deployment.
- Strong C++ and Python.
- Hands-on with embedded inference runtimes — LiteRT/TFLite, ONNX Runtime, ExecuTorch, TVM or vendor NPU SDKs — including delegate/execution-provider integration.
- Practical quantization expertise — INT8/INT4, per-channel schemes, calibration, accuracy recovery.
- Model formats and conversion tooling across PyTorch/TensorFlow to deployable graphs.
- Profiling on heterogeneous SoCs and reasoning about memory bandwidth as the dominant constraint.
- Understanding of CNN, transformer and contemporary vision/language model architectures.
Good-to-Have Skills
- On-device LLM/VLM deployment — llama.cpp class runtimes, speculative decoding, paged KV cache.
- Custom operator or kernel development for DSP/NPU/GPU (OpenCL, Vulkan compute, DSP intrinsics).
- Compiler-level work — MLIR, TVM, graph-level optimization passes.
- Vision pipeline integration and camera-to-inference zero-copy paths.
- MLOps for edge — model versioning, A/B evaluation, field accuracy monitoring.
📌 Technical Lead (Bengaluru)
🏢 HCLTech
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.