Overview
Build data systems that improve the quality, scale, and safety of data generated and acquired for frontier model training, spanning synthetic data generation and data anonymization/PII removal. Partner with researchers, domain experts, legal and compliance stakeholders, and customers to turn ambiguous data questions into experiments, pipelines, and durable products.
What you'll do
- Design and build systems that improve the quality, scale, and safety of data for frontier model training, spanning synthetic data generation and data anonymization/PII removal.
- Translate ambiguous research, partner, or compliance needs into clear hypotheses, experiments, evaluation plans, and production-quality implementations.
- Build and improve data-processing pipelines, evaluation frameworks, benchmarks, and quality-control systems.
- Run fast, rigorous iteration loops by prototyping, evaluating, interpreting results, and turning learnings into the next system or product.
- Partner with researchers, domain experts, and, where relevant, legal and compliance teams to ensure data is both high-utility and responsibly handled.
- Identify repeatable patterns across engagements and productize them into reusable software and platforms.
- Raise the technical bar through strong design judgment, clear communication, code quality, and mentorship.
What you'll need
- 2–10 years of recent, demonstrated experience in one or more of: synthetic/LLM-generated data, post-training and model-evaluation work, privacy engineering, or data anonymization/de-identification at scale.
- A hands-on individual contributor track record; this is not a team-lead or engineering-management role.
- Strong Python skills.
- Ability to write clean, efficient, scalable software for large, messy, real-world datasets.
- Sound judgment for reasoning about data quality, risk, and utility, including forming hypotheses, choosing meaningful metrics, diagnosing failures, and distinguishing signal from noise.
- Experience designing systems, including tradeoffs around quality, scale, reliability, and reuse.
- Comfort operating in an ambiguous, fast-moving environment with substantial ownership.
- Collaborative, low-ego communication and the ability to work effectively with researchers, engineers, domain experts, and customers.
Nice to have
- Experience building or operating large-scale synthetic or LLM-generated data pipelines for model training.
- Experience building or operating large-scale data de-identification or anonymization systems, ideally involving relational or graph-structured data, with experience preserving referential/relationship integrity after anonymization.
- Experience developing LLM/agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks.
- Research or applied work on reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems.
- Prior work in a regulated or high-sensitivity data environment, including healthcare, finance, HR/people data, or government.
- Experience with re-identification risk assessment and privacy auditing.
- Published research, meaningful open-source contributions, or evidence of technical leadership in ML systems, data engineering, or AI research.
- Experience productizing research or repeated customer work into robust, reusable platforms.
Details
- Location: Bangalore, India.
Read the full description and apply on the company’s own careers page.