Anthropic logo

Research Engineer / Scientist, Alignment

Anthropic
Posted a month ago

LOCATION

San Francisco · Hybrid

EXPERIENCE

Not specified

SALARY

$350K - $500K /year

SKILLS REQUIRED

Technical AI SafetyMachine Learning ExperimentsPythonLanguage Model Evaluation

Job description

Overview

Research Engineer / Scientist focused on Alignment Science, running exploratory AI safety experiments to understand and help steer risks from powerful future AI systems.

What you'll do

  • Build and run machine learning experiments related to AI safety and alignment.
  • Conduct exploratory experimental research on AI safety with a focus on risks from powerful future systems.
  • Collaborate with teams including Interpretability, Fine-Tuning, and the Frontier Red Team.
  • Train language models to subvert safety techniques and evaluate robustness.
  • Run multi-agent reinforcement learning experiments such as AI Debate.
  • Build tooling to evaluate the effectiveness of LLM-generated jailbreaks.
  • Write scripts and prompts to generate evaluation questions for safety-relevant reasoning.
  • Contribute materials to papers, blog posts, and talks.

What you'll need

  • Significant software, ML, or research engineering experience.
  • Experience contributing to empirical AI research projects.
  • Some familiarity with technical AI safety research.
  • Collaboration interest in fast-moving projects over extensive solo work.
  • Care about the impacts of AI.

Details

  • Interviews conducted in Python.
  • Preferred candidate location: based in the Bay Area.
  • Location-based hybrid policy: in offices at least 25% of the time.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Research Engineer / Scientist, Alignment