Overview
Research Engineer for Model Evaluations, designing and running evaluations to measure what Claude can do and to support reliable model-health monitoring during training.
What you'll do
- Design and run new evaluations of Claude capabilities (reasoning, agentic behavior, knowledge, safety properties) and visualize results.
- Build and harden a distributed evaluation execution platform to run hundreds of evals against checkpoints during production RL training runs.
- Own dashboards used by researchers and leadership to monitor model health (signal-to-noise, latency, regression detection).
- Debug anomalous eval results during training runs and determine whether the cause is model change or infrastructure issue.
- Improve tooling, libraries, and workflows for implementing and iterating on evaluations.
- Partner with research teams across defining what to measure through interpreting results as training progresses.
- Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks.
- Communicate evaluation results to internal stakeholders and, where appropriate, external audiences.
What you'll need
- Strong Python skills, including production or research infrastructure experience.
- Experience building or operating distributed systems and reliable-at-scale data pipelines or infrastructure.
- Clear written and verbal communication, especially when explaining technical results to non-specialists.
- Ability to operate in an on-call or production-support capacity while training runs are live.
- Care about societal impacts and interest in steering powerful AI to be safe and beneficial.
Details
- Annual compensation range: $500,000–$850,000 USD.
- Work mode: location-based hybrid policy; expected in an office at least 25% of the time.
- Visa sponsorship: provided, with reasonable efforts to get you a visa if you receive an offer.
Read the full description and apply on the company’s own careers page.