Overview
As a Performance Engineer, you will solve large-scale ML systems problems and build systems that optimize throughput and robustness for distributed systems.
What you'll do
- Identify novel systems problems that arise when running ML algorithms at scale.
- Develop systems to optimize throughput of large distributed systems.
- Improve robustness of distributed systems used for ML workloads.
- Work on large-scale systems problems with a focus on performance.
- Participate in pair programming.
- Design and implement performance and reliability improvements for production-like workloads.
What you'll need
- Significant software engineering or machine learning experience, particularly at supercomputing scale.
- Results-oriented approach with a bias towards flexibility and impact.
- Ability to take on tasks beyond the job description.
- Comfort with pair programming.
- Interest in learning more about machine learning research.
Nice to have
- Experience with high performance, large-scale ML systems.
- Experience with GPU/accelerator programming.
- Experience with ML framework internals.
- Experience with OS/kernel internals.
- Experience with language modeling with transformers.
Details
- Annual salary range: $280,000–$850,000 USD.
- Location-based hybrid policy: expected to be in an office at least 25% of the time.
- Visa sponsorship is available, with variability by role/candidate.
Read the full description and apply on the company’s own careers page.