Overview
Safeguards Enforcement Analyst on the User Well-being team, focused on designing and running mental health guardrails for AI users.
What you'll do
- Design and execute interventions, defining key metrics and curating evaluation datasets.
- Partner with Engineering and Data Science to build, tune, and validate detection models for automated intervention systems.
- Monitor how interventions and detection systems perform over time.
- Review flagged content to drive enforcement and policy improvements.
- Support in-product features that connect users to crisis resources and referral pathways.
- Provide feedback on policy gaps by analyzing real scenarios.
What you'll need
- Experience in trust & safety, product policy, content moderation, or a related field with exposure to mental health or related harms.
- Experience designing or running experiments, evaluations, or measurement studies for intervention effectiveness.
- Experience translating policy definitions into measurable form such as rubrics or classification criteria.
- Experience managing or coordinating content review operations, including quality assurance and workflow management.
- Proficiency in SQL and/or other data analysis tools to measure efficacy and monitor workflow health.
- Experience working with generative AI products, including writing prompts for content review and evaluation.
- Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
Details
- Compensation: annual salary range of $245,000 to $285,000 USD.
- Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.
- Location: hybrid policy requiring staff to be in an office at least 25% of the time.
- Visa sponsorship: sponsored visas are available for this role, subject to eligibility.
- Minimum years of experience: required years correlate with internal job level requirements (not explicitly stated).
Read the full description and apply on the company’s own careers page.