Overview
HPC Infrastructure Engineer to operate and improve ElevenLabs’ NVIDIA GPU clusters that power AI model training.
What you'll do
- Operate and improve the GPU fleet end to end, including provisioning, scheduling, monitoring, upgrades, and capacity planning.
- Build automation for node health checks, automated draining/remediation, and burn-in pipelines for new capacity.
- Own the stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and high-speed networking (InfiniBand/RoCE).
- Run and tune job scheduling using Slurm or similar.
- Build and maintain high-performance storage for datasets and checkpoints.
- Troubleshoot performance problems like stragglers, degraded links, thermal issues, and flaky GPUs.
- Keep clusters secure with access control, network isolation, and secrets.
What you'll need
- Experience operating large-scale Linux server or GPU environments in production.
- Knowledge of the NVIDIA stack (drivers, CUDA, NCCL, DCGM) or deep systems experience.
- Comfort with bare-metal environments, server hardware, and high-speed networking.
- Ability to write automation in Python and/or Bash, using IaC tools like Ansible or Terraform.
- Comfort analyzing metrics/logs/PromQL to find underlying issues.
- Ability to own scope end to end and be relied on by others.
- Willingness to support datacenter trips.
Nice to have
- Experience supporting ML training workloads from the infra side.
- Experience evaluating and working with GPU cloud providers.
- Experience with parallel filesystems or large-scale object storage.
- Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale.
- Power and cooling awareness for dense GPU deployments.
Details
- Remote role executable globally, with option to work from offices in London, New York, San Francisco, and Warsaw.
Read the full description and apply on the company’s own careers page.