Overview
As a Senior Data Engineer, design and maintain modern data architectures and large-scale data pipelines that support analytics, AI and GenAI use cases, and low-latency inference workloads.
What you'll do
- Design, build, and maintain robust ETL/ELT pipelines supporting large, complex datasets with evolving schemas.
- Support embedding generation and vector-based data pipelines for AI and GenAI use cases.
- Develop scalable batch and streaming data pipelines using cloud-based data platforms such as Snowflake or Databricks.
- Create and evolve semantic data models that transform raw data into analytics-ready, trustworthy datasets.
- Build data preprocessing, validation, and quality-assurance tooling to ensure reliability and correctness.
- Design pipelines that support both offline model training and low-latency inference workloads.
- Drive the design of modern data architectures that balance scalability, performance, and maintainability.
- Implement data quality checks, lineage, and monitoring for enterprise-scale datasets.
- Work closely with product, analytics, AI/ML, IT, and business teams to understand data needs and enable data-driven initiatives.
- Interface with stakeholders across manufacturing, quality, post-market surveillance, commercial, and clinical domains.
- Provide technical leadership and mentorship to other data engineers, promoting best practices and continuous improvement.
What you'll need
- Bachelor’s Degree or higher in Computer Science, Mathematics, engineering or a related technical field and 6-8 years of job related experience, or High School Diploma/GED and 10 years of the same job-related experience.
- Strong experience with SQL, relational and NoSQL databases.
- Hands-on experience building and maintaining large-scale data pipelines in cloud environments using Azure or AWS.
- Experience with Databricks or Snowflake and distributed data processing frameworks such as Spark.
- Familiarity with feature stores, vector databases, or model-adjacent data systems.
- Experience managing data quality, validation, and lineage for large, complex datasets.
- Proficiency in Python and experience with data-frame libraries such as Pandas or Polars.
- Experience with ETL/workflow orchestration tools such as Databricks Workflows, Azure Data Factory, or AWS Glue.
- Familiarity with common data formats, including CSV, JSON, XML, and Parquet.
- Strong communication skills with the ability to explain technical concepts to non-technical stakeholders.
- Professional proficiency in English, with regular collaboration across the global team.
Nice to have
- Experience in healthcare, medical device, life sciences, or manufacturing environments.
- Familiarity with Business Intelligence tools and analytics platforms.
- Experience supporting AI/ML use cases through high-quality, well-modeled data.
- Knowledge of data privacy, security, and compliance best practices.
Details
- Work location: Hybrid.
- Domestic travel may include up to 20%.
- You may be required to enter healthcare or other third-party facilities, which may require certain licenses, vaccinations, and/or other credentials or qualifications.
Read the full description and apply on the company’s own careers page.