Overview
Sr Machine Learning Engineer for Amgen’s Manufacturing Applications Product Team, focused on designing and optimizing data pipelines and ingestion solutions that enable AI/LLM-powered knowledge layer and assistant experiences.
What you'll do
- Design, develop, and maintain complex ETL/ELT data pipelines in Databricks using PySpark, Scala, and SQL.
- Build ML/LLM feature engineering for ingestion, embeddings/vector DBs, RAG/LLM serving, and latency optimization.
- Implement data access, logging, privacy controls, governance, and interoperability across hybrid cloud environments.
- Ingest and transform structured and unstructured data from databases, APIs, logs, event streams, and third-party platforms.
- Ensure data integrity, accuracy, and consistency using quality checks and monitoring.
- Collaborate in Agile/Scaled Agile (SAFe) delivery with cross-functional teams and manage work with JIRA/Confluence.
What you'll need
- Hands-on data engineering experience with Databricks, PySpark/SparkSQL, Apache Spark, AWS, Python, SQL, and SAFe.
- Workflow orchestration and performance tuning experience for big data processing.
- Strong understanding of AWS services.
- Experience with streaming technologies such as Apache Kafka or Debezium.
- Experience collaborating using Scaled Agile Framework (SAFe) and Agile/DevOps practices.
- Doctorate degree OR Master’s with 4–6 years OR Bachelor’s with 6–8 years in Computer Science/IT or related field.
Nice to have
- Experience with AI-assisted code development tools like GitHub Copilot, Cursor, or Claude Code.
- Experience with biotech/pharma/manufacturing data sources (e.g., SCADA, Data Historian) and industry collaboration.
- Experience writing APIs for data consumers.
- Experience with vector databases, data modeling, and performance tuning for OLAP and OLTP.
Details
- Works in an Agile and Scaled Agile (SAFe) environment.
- Uses JIRA and Confluence to manage sprints, backlogs, and user stories.