Job DescriptionAIML Pipeline / Dataflow Design Lead
nLocation: Bangalore, India |Company: LTSCT — L&T; Semiconductor Technologies Limited
nReporting To: Senior Architect, DC HW Systems | Experience: 10+ years (mandatory)
nRole Summary
nWe are hiring two AIML Pipeline / Dataflow Design Leads to implement the BFOS-MPE (Broadcast-Fed Output-Stationary Matrix Processing Engine) — the compute heart of xPU. You will own the compute datapath, the broadcast network, PE array control, sparsity engines (SWS), and dual-mode (matrix + GPU) switching. You will also own the LLVM interface to the xPU.
nKey Responsibilities
n
- n
- Own micro-architecture and RTL of the BFOS-MPE compute datapath — dual 256×256 PE arrays (65,536 PEs each) operating at 2 GHz.n
- Design the broadcast-fed input network and output-stationary accumulation scheme that define the BFOS dataflow.n
- Implement PE array control, sequencing, and the sparsity engine (SWS) for structured/unstructured sparsity acceleration.n
- Architect dual-mode operation enabling switching between dense matrix (GEMM/convolution)
and GPU-style workloads.n
- Drive numerical formats and precision (FP/BF/INT, mixed precision) and validate functional correctness against reference models.n
- Coordinate and integrate contributions into the dataflow codebase and verification plan.n
- Own performance/utilization modeling, power-efficiency optimization, and PPA closure for the compute engine.n
- Own the ISA/FW routines for the optimal utilization of the HW compute structures in the xPU; as part of the LLVM development.n
nRequired Skills & Experience
n
- n
- 10+ years in compute datapath, DSP, GPU, or AI-accelerator micro-architecture and RTL design.n
- Deep understanding of systolic/spatial arrays, matrix-multiply dataflows (output/weight/row-stationary), and broadcast networks.n
- Solid expertise in RTL (SystemVerilog) for high-throughput arithmetic datapaths and control.n
- Solid grounding in numerical formats and a