31 Jul
|
PowaiILabs Technology
|
Mumbai
31 Jul
PowaiILabs Technology
Mumbai
About the Role We are looking for a rare, high-impact engineer who operates at the intersection of Deep Learning Architecture and Hardware Design. In the initial phase, you will focus on implementing, benchmarking, and optimizing state-of-the-art Computer Vision models—starting with ResNet variants, YOLO, UNet, DenseNet, ConvNeXt, MobileNet (v1–v4), and custom deep networks. As our product evolves, you will transition into Hardware-Software Co-Design, leading the effort to map these neural network workloads onto custom hardware (FPGA / ASIC / NPU).
You will bridges the gap between high-level PyTorch models and low-level hardware constraints (compute bandwidth, memory hierarchy, SRAM limits, and quantization).
Key Responsibilities Phase 1: Deep Learning Workloads & Neural Architecture (Immediate) Design, train, optimize, and benchmark Computer Vision and CNN architectures, including ResNet, YOLO, UNet, Wide ResNet, DenseNet, Inception-ResNet, ConvNeXt, MobileNet (v1–v4), and specialized deep networks.
Perform model profiling to identify compute bottlenecks (FLOPs vs.
Memory
Bandwidth, MAC intensity, layer latency).
Implement model compression techniques tailored for downstream hardware target execution: Quantization (INT8/FP8/INT4), Pruning, Knowledge Distillation, and Channel Sparsity.
Convert models to deployment runtimes (ONNX, TensorRT, TVM, TFLite) and benchmark latency/power constraints.
Phase 2: Hardware Architecture & Silicon Co-Design (Future Expansion) Act as the lead link between AI algorithm design and custom silicon/FPGA implementation.
Define hardware compute requirements for neural processing units (NPUs/accelerators) based on target model operations (2D Convolutions, Depthwise Separable Convolutions, Residual Connections, Feature Map Concatenation).
Design, simulate,
and model hardware execution units in SystemVerilog/Verilog/VHDL or high-level modeling frameworks (C++/SystemC). Optimize memory access patterns (SRAM buffering, line buffers, DMA transfers, and DDR/HBM bandwidth) for CNN features.
Work with EDA toolchains for synthesis, timing analysis, and FPGA target prototyping.
Required Qualifications & Skills 1.
AI & Deep Learning Software Frameworks: Expert proficiency in PyTorch or TensorFlow.
Model Expertise: Hands-on experience training, modifying, and exporting CNN architectures:
Classification/Backbones: ResNet, Wide ResNet, DenseNet, Inception-ResNet, ConvNeXt, MobileNet (v1/v2/v3/v4).
Detection & Segmentation: YOLO (v5/v8/v10/v11), UNet.
Optimization: Deep understanding of layer-by-layer execution, quantization-aware training (QAT), weight pruning, and operator fusion.
2.
Hardware & Architecture Experience RTL / HDLs: Proficiency in SystemVerilog, Verilog, or VHDL.
Hardware Architecture: Understanding of Computer Architecture fundamentals: MAC units, Systolic Arrays, Memory Hierarchies (SRAM/DRAM caches), Dataflow Architectures (Weight Stationary vs.
Output
Stationary), and AXI/AMBA bus protocols.
Prototyping: Experience targeting FPGAs (Xilinx Vivado, Intel Quartus) or custom ASIC methodologies for deep learning acceleration.
Programming Languages: High proficiency in Python and C / C++. Preferred / Nice-to-Have Qualifications Experience with AI compiler stacks like TVM, MLIR, ONNX Runtime, or Glow.
Prior participation in an ASIC Tapeout or FPGA-based deep learning accelerator deployment.
Familiarity with low-power edge platforms (NVIDIA Jetson, NXP, Hailo, ESP32-S3, or STM32 MCU AI toolchains).
Master's or Ph.D. in Electrical Engineering, Computer Engineering, or Computer Science with a focus on Hardware Acceleration for AI.
📌 AI Hardware & Deep Learning Engineer (Edge / ASIC / FPGA) (Mumbai)
🏢 PowaiILabs Technology
📍 Mumbai