Senior Technical Architect (Hyderabad)

Senior Technical Architect (Hyderabad)

20 Aug
|
HCLTech
|
Hyderabad

20 Aug

HCLTech

Hyderabad

Hyderabad, Telangana

Job Summary

Own the on-device AI execution stack for ARM-based SoCs with dedicated neural accelerators — model conversion and quantization, runtime and delegate integration, and deployment of vision, LLM and VLM workloads within edge power and memory budgets.

Key Responsibilities

Deploy and optimize neural network models on NPU, GPU and CPU backends using embedded inference runtimes (LiteRT/TensorFlow Lite, ONNX Runtime, ExecuTorch or vendor SDKs).

Own the model conversion pipeline — graph capture, operator mapping, quantization (PTQ and QAT support), calibration and accuracy validation against reference.

Diagnose and resolve unsupported operators, graph partitioning and fallback behaviour; work with vendor toolchains on accelerator limitations.

Enable on-device LLM and VLM workloads — weight quantization, KV-cache management, prefill/decode optimization, memory footprint and token-throughput tuning.

Benchmark and optimize inference latency, throughput, memory bandwidth and energy per inference; publish reproducible performance data.





Integrate inference into real-time media and robotics pipelines with zero-copy tensor/buffer sharing.

Build model deployment and regression tooling so accuracy and performance are tracked across releases.

Advise product and data-science teams on model architecture choices suitable for edge accelerators.

Skill Requirements

8-10 years in software engineering, with 3+ years in on-device / edge AI deployment.

Solid C++ and Python.

Hands-on with embedded inference runtimes — LiteRT/TFLite, ONNX Runtime, ExecuTorch, TVM or vendor NPU SDKs — including delegate/execution-provider integration.

Practical quantization expertise — INT8/INT4, per-channel schemes, calibration, accuracy recovery.

Model formats and conversion tooling across PyTorch/TensorFlow to deployable graphs.

Profiling on heterogeneous SoCs and reasoning about memory bandwidth as the dominant constraint.

Understanding of CNN, transformer an

📌 Senior Technical Architect (Hyderabad)
🏢 HCLTech
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior technical architect (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: senior technical architect (hyderabad) / hyderabad