Overview
Build and run cyber-focused evaluations that measure model capabilities, safeguard robustness, and cyber-abuse risks. The role also develops detection probes and supports abuse-detection architecture and safeguard improvements.
What you'll do
- Design and run capability, uplift, and safety evaluations for cyber-relevant risks.
- Execute safeguard-robustness testing before major model launches.
- Analyze evaluation results and communicate findings to stakeholders.
- Design, prototype, and tune cyber-misuse detection probes.
- Translate policy requirements into layered abuse-detection architecture.
- Build and maintain tooling for running and scoring evaluations.
- Collaborate on translating evaluation findings into safeguard improvements.
What you'll need
- Experience building or running evaluations, benchmarks, or test suites for software or ML systems.
- Hands-on cybersecurity experience.
- Proficiency in Python.
- Strong communication skills with cross-functional and policy stakeholders.
- Bachelor’s degree or equivalent education, training, and experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
Nice to have
- Deep offensive-security or security-research experience.
- Experience analyzing adversarial or abuse data.
- Experience with AI/ML evaluation frameworks.
- Experience authoring detection content or building ML-based abuse detection.
- Active secret security clearance or eligibility to obtain one.
Details
- Annual salary: $300,000–$405,000 USD.
- Hybrid policy: staff are expected in an office at least 25% of the time.
- Locations include San Francisco, California, and Washington, DC, United States.
Read the full description and apply on the company’s own careers page.