Senior Engineer - AI Inference & Kernel Optimization (India)

Senior Engineer - AI Inference & Kernel Optimization (India)

16 Aug
|
SQUARE ROOT CONSULTING
|
India

16 Aug

SQUARE ROOT CONSULTING

India

Description :

Position : Sr Engineer AI Inference & Kernel Optimization

Location : Bangalore, India (On-site / Hybrid)

Compensation : ?7090 LPA (Base) ESOPs Performance Variable

(Compensation aligned with experience and technical depth)

About the Role :

We are building next-generation AI inference processors optimized for ultra-low latency, high-throughput workloads. As a Senior / Principal Engineer, you will play a critical role in designing and optimizing low-level software and compute kernels that extract maximum performance from GPUs and custom accelerators.

This role is ideal for engineers who thrive close to hardware, enjoy performance tuning, and want to influence the software stack of cutting-edge AI silicon.

Key Responsibilities :

- Design, implement, and optimize high-performance compute kernels for GPUs and/or custom AI accelerators
- Develop low-level software in C/C and CUDA, targeting inference workloads for deep learning models
- Apply advanced code optimization techniques, including :

a. Vectorization (SIMD)

b. Memory hierarchy optimization (registers, shared memory, caches)

c. Parallelization strategies

- Cache utilization and memory bandwidth optimization
- Drive profiling, benchmarking, and performance tuning to achieve optimal resource utilization and minimal inference latency
- Collaborate closely with architecture, compiler, and hardware teams to co-design performant solutions
- Analyze bottlenecks across compute, memory, and interconnects, and propose architectural or software improvements
- Mentor junior engineers and contribute to technical direction (for Principal-level candidates)





Required Qualifications :

- Strong expertise in C and C , with deep understanding of low-level programming
- Hands-on experience with CUDA and GPU programming
- Proven experience developing high-performance kernels
- Deep knowledge of performance optimization techniques, including :

a. Vectorization and instruction-level optimization b. Threading and parallel execution models c. Memory hierarchies and cache behavior

- Experience with profiling and performance analysis tools (e.g., Nsight, VTune, perf, custom profilers)
- Strong understanding of AI inference workloads (CNNs, Transformers, GEMM, attention, activation functions, etc.)

Preferred / Nice-to-Have :

- Experience working on AI inference frameworks, runtimes, or compilers
- Background in computer architecture or microarchitecture
- Experience optimizing for latency-critical systems
- Exposure to custom silicon bring-up or hardware-software co-design
- Contributions to performance-critical open-source projects

Level Expectations :

Senior Engineer :

- Own complex kernel implementations
- Independently drive optimization and tuning
- Deliver production-quality, high-performance code

Principal Engineer :

- Define performance strategy and best practices
- Influence architecture and software stack decisions
- Lead complex, cross-functional technical initiatives

Why Join Us :

- Work on cutting-edge AI inference silicon
- Own performance-critical parts of the stack, close to hardware
- High-impact role with solid technical ownership
- Competitive compensation with meaningful ESOP upside

📌 Senior Engineer - AI Inference & Kernel Optimization (India)
🏢 SQUARE ROOT CONSULTING
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior engineer - ai inference & kernel optimization (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior engineer - ai inference & kernel optimization (india) / india