Overview
Serve as a Principal Engineer for scale-up GPU networking in high-performance computing and AI, delivering high-bandwidth, low-latency intra-node communication and GPU-accelerated communication workflows.
What you'll do
- Architect and deliver GPU-aware networking paths for high-bandwidth, low-latency intra-node communication.
- Develop and optimize GPU → NIC → GPU data movement, shared memory models, and DMA pathways.
- Work with NVIDIA CUDA, NVLink, NCCL, AMD ROCm, InfinityFabric, and RCCL teams to integrate and optimize scale-up communication semantics.
- Drive improvements to DMA engines, BAR mappings, ATS/IOMMU, and GPU memory registration workflows.
- Enhance and extend Libfabric, UCX, CXI, SHMEMX, and OpenMPI for GPU-accelerated scale-up workflows.
- Optimize communication collectives, transport layers, and GPU-direct capabilities.
- Characterize and tune multi-NIC per socket, NUMA-zone mapping, GPU locality, CQ/queue design, and CPU/GPU topology optimization.
- Lead upstream contributions to open-source projects including OFI, UCX, OpenMPI, and RCCL/NCCL enablement.
- Partner with HPC/AI ecosystem teams to shape future architectures.
- Own complex debugging across driver, runtime, GPU, kernel, and user-space boundaries.
- Develop profiling workflows using Nsight, ROCm tools, eBPF, perf, and other tools.
What you'll need
- 10–15+ years building high-performance networking, GPU, or kernel-level software.
- Deep expertise in C/C++, Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, and DMA engines.
- Strong understanding of CUDA, ROCm, GPU memory models, P2P, GDS (GPUDirect Storage), and GDR (GPUDirect RDMA).
- Hands-on experience with MPI, SHMEM, Libfabric, UCX, or similar communication stacks.
- Proven experience driving architecture, cross-org technical decisions, and upstream contributions.
- Ability to mentor senior engineers, influence multi-team designs, and own end-to-end delivery.
Nice to have
- Experience with NIC architecture including CXI, RoCE, Infiniband, Slingshot, and NVLink Switch.
- Experience optimizing collectives such as AllReduce and AllGather on GPUs.
- Background contributing to open-source HPC/AI libraries.
- Familiarity with HPC system architecture, NUMA tuning, and multi-accelerator systems.
Details
- Location: Bengaluru, Karnataka, India.
- Hybrid role requiring an average of 2 days per week from an HPE office.
Read the full description and apply on the company’s own careers page.