Overview
Senior Data Scientist role focused on characterizing conversational audio data and building data strategies that improve speech and language models. Collaborate with research, engineering, and data teams on rigorous, scalable experiments and tooling.
What you'll do
- Characterize languages, conditions, domains, speakers, quality, and underrepresented areas in data.
- Design active-learning loops to prioritize work based on expected performance impact.
- Design human-in-the-loop workflows, tooling, and model-assisted processes.
- Build curated datasets and benchmarking methods for representative speech evaluation.
- Run experiments to determine which data strategies improve model performance.
- Build repeatable pipelines for domain- and customer-specific model adaptation.
What you'll need
- Hands-on experience with data pipelines and model-facing problems in data science, machine learning, or applied research.
- Strong Python and data-tooling skills.
- Experience with data characterization, data selection, active learning, or similar prioritization problems.
- Familiarity with speech/audio or NLP models, including output quality, confidence, and error modes.
- Experience turning ambiguous data problems into measurable model or product improvements.
- Ability to build reusable systems and communicate complex findings to varied audiences.
- Active use of AI tools in work.
Nice to have
- Experience with ASR/TTS, audio data, or multilingual and code-switched data.
- Experience with ensemble labeling, pseudo-labeling, or LLM-assisted annotation.
- Familiarity with data provenance, PII/GDPR-aware pipelines, or model-improvement compliance.
- Experience building custom or fine-tuned models for customers or domains.
- Experience working with research and engineering teams on shared infrastructure.
Details
Read the full description and apply on the company’s own careers page.