Overview
Staff Software Engineer in AI Reliability Engineering (AIRE) to improve the reliability of Claude across the full serving path—from SDK through networks, API layers, and infrastructure to accelerators and back.
What you'll do
- Develop service level objectives (SLOs) for large language model serving while balancing availability, latency, and development velocity.
- Design and implement monitoring and observability along the token path.
- Design and implement high-availability serving infrastructure across multiple regions and cloud providers.
- Lead incident response for critical AI services, including recovery, incident reviews, and systematic improvements.
- Support reliability of safeguard model serving for both site reliability and safety commitments.
What you'll need
- Strong distributed systems, infrastructure, or reliability background.
- Comfort jumping into unfamiliar systems during incidents and driving resolution.
- Ability to think holistically about how systems compose and where interfaces/seams are.
- Ability to build lasting relationships across teams.
- Excellent communication and collaboration skills.
- Care about users and ownership over outcomes.
Details
- Annual compensation range: £325,000 to £390,000 GBP.
- Logistics: expects staff to be in one of the offices at least 25% of the time (hybrid policy).
- Visa sponsorship is available in some cases, with reasonable efforts made if an offer is made.
- Minimum education: Bachelor’s degree or equivalent education/training/experience combination.
- Required field of study: field relevant to the role as demonstrated by coursework, training, or professional experience.
Read the full description and apply on the company’s own careers page.