02 Sep
|
L&T Semiconductor Technologies
|
Bengaluru
02 Sep
L&T Semiconductor Technologies
Bengaluru
Job Description AIML Pipeline / Dataflow Design Lead N Location: Bangalore, India |Company: LTSCT — L&T; Semiconductor Technologies Limited N Reporting To: Senior Architect, DC HW Systems | Experience: 10+ years (mandatory) N Role Summary nWe are hiring two AIML Pipeline / Dataflow Design Leads to implement the BFOS-MPE (Broadcast-Fed Output-Stationary Matrix Processing Engine) — the compute heart of xPU. You will own the compute datapath, the broadcast network, PE array control, sparsity engines (SWS), and dual-mode (matrix + GPU) switching. You will also own the LLVM interface to the xPU. N Key Responsibilities Nn Own micro-architecture and RTL of the BFOS-MPE compute datapath — dual 256×256 PE arrays (65,536 PEs each) operating at 2 GHz. N Design the broadcast-fed input network and output-stationary accumulation scheme that define the BFOS dataflow. N Implement PE array control, sequencing, and the sparsity engine (SWS) for structured/unstructured sparsity acceleration. N Architect dual-mode operation enabling switching between dense matrix (GEMM/convolution) and GPU-style workloads. N Drive numerical formatsand precision (FP/BF/INT, mixed precision) and validate functional correctness against reference models. N Coordinate and integrate contributions into the dataflow codebase and verification plan. N Own performance/utilization modeling, power-efficiency optimization, and PPA closure for the compute engine.
N Own the ISA/FW routinesfor the optimal utilization of the HW compute structures in the xPU; as part of the LLVM development. Nn Required Skills & Experience Nn 10+ years in compute datapath, DSP, GPU, or AI-accelerator micro-architecture and RTL design. N Deep understanding of systolic/spatial arrays, matrix-multiply dataflows (output/weight/row-stationary), and broadcast networks. N Solid expertise in RTL(SystemVerilog) for high-throughput arithmetic datapaths and control. N Solid grounding in numerical formats and arithmetic (FP32/BF16/FP8/INT8), mixed-precision, and rounding/accuracy trade-offs. N Experience with sparsity acceleration techniques (structured/unstructured, weight/activation sparsity). N Familiarity with deep-learning operators (GEMM, convolution, attention) and how they map to hardware. N Ability to build and correlate performance/utilization models against RTL and reference software. N Hands-on PPA optimization for large, dense compute blocks. N Proficiency in C/C++/Python for modeling and verification support. N Proven leadership and experience integrating cross-organization/partner engineering contributions. N AI compiler developmentexperience is a definite plus. Nn Preferred Qualifications Nn Direct experienceon a taped-out AI/ML accelerator or tensor/matrix engine. N Familiarity with GPU SIMT execution models and matrix+vector dual-mode designs. N Exposure to ML compiler/graph mapping (MLIR, TVM) and how software drives the datapath. N M.Tech/MS/PhD in EE/ECE/CS or equivalent. N
📌 Aiml Pipeline/ Dataflow Design Lead (Bengaluru)
🏢 L&T Semiconductor Technologies
📍 Bengaluru