Overview
Software Engineer on the ML Infrastructure team, designing and building platforms for scalable, reliable, and efficient serving of LLMs.
What you'll do
- Build and maintain fault-tolerant, high-performance systems for serving LLM workloads at scale.
- Build an internal platform to enable LLM capability discovery.
- Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
- Conduct architecture and design reviews to uphold best practices in system design and scalability.
- Develop monitoring and observability solutions for system health and performance.
- Lead projects end-to-end from requirements gathering to implementation in a cross-functional environment.
What you'll need
- 4+ years of experience building large-scale, high-performance backend systems.
- Strong programming skills in one or more languages such as Python, Go, Rust, or C++.
- Experience with LLM serving and routing fundamentals including rate limiting, token streaming, load balancing, and budgets.
- Experience with LLM capabilities and concepts such as reasoning, tool calling, and prompt templates.
- Experience with containers and orchestration tools such as Docker and Kubernetes.
- Familiarity with cloud infrastructure such as AWS and GCP, and infrastructure as code such as Terraform.
- Proven ability to solve complex problems and work independently in fast-moving environments.
Nice to have
- Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.
Details
- Full-time role; base salary range stated for San Francisco, New York, and Seattle.
Read the full description and apply on the company’s own careers page.