Overview
AI and HPC Systems Performance Engineer focused on tuning GPU server performance for AI training and inference workloads on Linux-based systems.
What you'll do
- Install, configure, and optimize AI infrastructure components including GPU servers, storage, networking, and AI software stacks.
- Capture, analyze, and interpret telemetry, logs, traces, and profiling data to characterize workload execution and find bottlenecks.
- Design, execute, and analyze performance benchmarks for AI/ML workloads including LLMs and RAG pipelines.
- Optimize performance across multi-GPU and distributed AI environments using interconnect, storage, and networking technologies.
- Develop automation scripts, deployment frameworks, and Infrastructure-as-Code solutions for AI platform provisioning.
- Collaborate with customers, partners, and internal engineering teams to troubleshoot and optimize AI solutions on HPE platforms.
What you'll need
- Typically 8+ years of experience.
- Strong Linux system administration and command-line experience across enterprise Linux distributions.
- Experience with modern AI/ML frameworks and ecosystems including PyTorch, JAX, and Hugging Face Transformers.
- Experience with AI training/inference, benchmarking, performance characterization, and optimization.
- Experience with high-performance networking technologies including InfiniBand and RDMA (incl. RoCE) and networking solutions.
- Proficiency in one or more programming/scripting languages such as Python, Bash, Go, C++, or similar.
- Experience with containerized/orchestrated environments including Docker and Kubernetes.
Details
- Hybrid role with an average requirement to work 2 days per week from an HPE office.
Read the full description and apply on the company’s own careers page.