Overview
You’ll lead operational health and forward momentum for Safeguards Engineering infrastructure and evaluation platforms, with a focus on reliability for safety-critical ML systems.
What you'll do
- Own the Safeguards Engineering operations review cadence and reliability trend visibility.
- Drive incident tracking and post-mortem execution, including completing resulting action items.
- Define and maintain SLOs for safety-critical pipelines with partner teams and track/report performance.
- Maintain runbook quality and ensure clear incident ownership during active incidents.
- Manage platform migrations and infrastructure projects across incident and monitoring platforms.
- Coordinate evaluation platform improvements, including self-serve capabilities and eval factory infrastructure.
- Triages issues by understanding how production ML systems work and what is safety-critical.
What you'll need
- Solid technical program management experience, including operational/infrastructure-heavy environments.
- Technical understanding of production ML systems to triage incidents and discuss issues with engineers.
- Comfort coordinating across boundaries where you may lack direct authority.
- The ability to balance “keep the lights on” operations with longer-horizon platform projects.
- Ability to build processes that close loops on post-mortem actions, SLO checks, and runbook upkeep.
- Interest in AI safety and understanding reliability of safety-critical pipelines as distinct from product reliability.
Details
- Hybrid work policy: expected in an office at least 25% of the time.
Read the full description and apply on the company’s own careers page.