O

Researcher, Alignment CoT Monitorability

OpenAI
Posted a month ago

LOCATION

San Francisco · Hybrid

EXPERIENCE

Not specified

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

Alignment MonitoringLLM EvaluationModel InterpretabilityEmpirical Machine Learning

Job description

Overview

Researcher on the CoT Monitorability/Alignment team studying when model chain-of-thought is monitorable enough for scalable oversight. You will design and run empirical experiments to improve understanding of model monitorability and translate findings into oversight and training recommendations.

What you'll do

  • Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings.
  • Build evaluations to measure whether monitors can reliably predict properties of interest, including high-stakes misbehavior.
  • Investigate how interventions across the training pipeline (pre-training, synthetic data, mid-training, post-training, reinforcement learning, others) improve or degrade monitorability.
  • Analyze model behavior and convert monitoring observations into hypotheses, experiments, and recommendations.
  • Translate research findings into practical monitoring and oversight approaches for real training runs.
  • Collaborate with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work.
  • Produce externally publishable research when results advance broader alignment science.

What you'll need

  • Strong hands-on empirical ML expertise.
  • Deep interest in model behavior, alignment, or interpretability.
  • Ability to train, evaluate, or debug large ML models, especially LLMs.
  • Depth in alignment, interpretability, model behavior, empirical ML, or adjacent research.
  • Strong skill in turning ambiguous model-behavior questions into measurable experiments.
  • Ability to move between research ideation and engineering execution.
  • High independence with close collaboration across research and engineering teams.

Details

  • Based in San Francisco, CA.
  • Hybrid work model: 3 days in the office per week.
  • Relocation assistance for new employees.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Researcher, Alignment CoT Monitorability