Anthropic logo

Research Engineer, Interpretability

Anthropic
Posted a month ago

LOCATION

San Francisco · Hybrid

EXPERIENCE

5 - 10 Years

SALARY

$315K - $560K /year

SKILLS REQUIRED

Model InterpretabilityNeural Network TrainingPerformance ProfilingMachine LearningPython

Job description

Overview

Research Engineer (Interpretability) for building infrastructure and tools that enable mechanistic interpretability of neural networks and support production safety audits.

What you'll do

  • Build and maintain specialized inference and training infrastructure for interpretability research.
  • Resolve scaling and efficiency bottlenecks via profiling and optimization.
  • Design tools, abstractions, and platforms to help researchers experiment rapidly.
  • Support bringing interpretability research into production safety audits with reliability.
  • Work across the stack from model internals to accelerator-level optimization and research tooling.

What you'll need

  • Have 5-10+ years of experience building software.
  • Be highly proficient in at least one programming language (e.g., Python, Rust, Go, Java) and productive with Python.
  • Demonstrate curiosity to learn unfamiliar domains quickly.
  • Be able to prioritize impactful work and operate comfortably with ambiguity.
  • Collaborate in fast-moving projects and translate research needs into engineering solutions.
  • Have interest in interpretability research and its role in AI safety (no formal research experience required).

Details

  • Based in the San Francisco office, with case-by-case consideration for remote work.
  • Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Research Engineer, Interpretability