Overview
The Senior Data Scientist - Protein Data Pipelines will enable predictive modeling for protein sequence, structure, and function by building scalable, reliable, and reproducible data pipelines for ML use.
What you'll do
- Design and maintain scalable data pipelines for predictive model training focused on protein sequence or structure-to-function applications.
- Build ML-amenable data assets from protein property data that are readable, quality-controlled, and reproducible for reuse.
- Develop deployment strategies and pipelines to embed trained models into ongoing projects and support inference workflows.
- Create reusable inference, deployment, and testing frameworks for in-house and external ML models.
- Establish data quality, validation, monitoring, and reproducibility practices for protein datasets and related outputs.
- Document data lineage, assumptions, validation outcomes, and reproducibility practices to support long-term reuse.
What you'll need
- Bachelor’s degree in computational biology/bioinformatics/life sciences/data science or a related field plus relevant professional experience.
- Meet one of: 6+ years (bachelor’s), 4+ years (master’s), or a PhD (experience requirement not specified beyond the degree condition).
- Strong experience building scalable data pipelines in Python and/or SQL.
- Hands-on experience owning reusable, end-to-end MLOps for at least one ML model.
- Knowledge of data quality control, validation, and monitoring practices.
- Familiarity with computational/related scientific fields, including protein sequence/structure/property datasets.
Details
- Location: Hyderabad, India.
Read the full description and apply on the company’s own careers page.