Principal Engineer - Scale-Up GPU Networking (HPC / AI) (Bengaluru)

Principal Engineer - Scale-Up GPU Networking (HPC / AI) (Bengaluru)

06 Oct
|
Hewlett Packard Enterprise
|
Bengaluru

06 Oct

Hewlett Packard Enterprise

Bengaluru

Principal Engineer - Scale-Up GPU Networking (HPC / AI)

This role has been designed as Hybrid with a requirement that you will work on average 2 days per week from an HPE office.

High Performance Computing, AI and Labs is a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers digital transformation, enabling them to tackle their complex, and data-intensive workloads. Combining deep expertise and the development of the world s most cutting-edge, high-performance supercomputers, is defining the next era of computing delivering valuable insight innovation. Join us and redefine what s next for you.

What youll do:
Key Responsibilities
Architect Deliver Scale-Up Networking

- Design and implement GPU-aware networking paths for high-bandwidth, low-latency intra-node communication.
- Develop and optimize GPU NIC GPU data movement, shared memory models, and DMA pathways.

GPU Ecosystem Integration

- Work with NVIDIA CUDA, NVLink, NCCL , and AMD ROCm, InfinityFabric, RCCL teams to integrate and optimize scale-up communication semantics.
- Drive improvements to DMA engines, BAR mappings, ATS/IOMMU , and GPU memory registration workflows.

Runtime Communication Stack Development

- Enhance and extend Libfabric, UCX, CXI, SHMEMX, OpenMPI for GPU-accelerated scale-up workflows.
- Optimize communication collectives, transport layers, and GPU-direct capabilities.

Multi-NIC / NUMA Performance Optimization

- Characterize and tune multi-NIC per socket , NUMA-zone mapping, GPU locality, CQ/queue design, and CPU/GPU topology optimization.

Upstreaming Architecture Influence

- Lead upstream contributions to open-source projects (OFI, UCX, OpenMPI, RCCL/NCCL enablement).
- Partner with HPC/AI ecosystem teams to shape future architectures.

Debugging, Performance, and Quality

- Own complex debugging across driver, runtime, GPU, kernel, and user-space boundaries.
- Develop profiling workflows using Nsight, ROCm tools, eBPF, perf, etc.

What you need to bring:
Required Skills Experience

- 10-15+ years building high-performance networking, GPU, or kernel-level software .
- Deep expertise in C/C++ , Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, DMA engines.
- Solid understanding of CUDA, ROCm,



GPU memory models, P2P, GDS (GPUDirect Storage), GDR (GPUDirect RDMA) .
- Hands-on experience with MPI, SHMEM, Libfabric, UCX , or similar communication stacks.
- Proven experience driving architecture , cross-org technical decisions, and upstream contributions.
- Ability to mentor senior engineers, influence multi-team designs, and own end-to-end delivery.

Preferred Qualifications

- Experience with NIC architecture (CXI, RoCE, Infiniband, Slingshot, NVLink Switch).
- Experience optimizing collectives (AllReduce/AllGather) on GPUs.
- Background contributing to open-source HPC/AI libraries .
- Familiarity with HPC system architecture , NUMA tuning, and multi-accelerator systems.

Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here .

Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:

Health Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs.



We make bold moves, together, and are a force for good.

Lets Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#india #highperformancecompute

Job:

Engineering

Job Level:

TCP_05

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity .

Recruitment Fraud Alert

We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.

All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual s own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Principal Engineer - Scale-Up GPU Networking (HPC / AI) (Bengaluru)
🏢 Hewlett Packard Enterprise
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer - scale-up gpu networking (hpc / ai) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer - scale-up gpu networking (hpc / ai) (bengaluru) / bengaluru