Overview
Principal Data Engineer - AI responsible for setting technical direction for ingesting, transforming, storing, serving, and governing data across the full stack of a data platform.
What you'll do
- Lead data architecture, design, and deployment of scalable, high-throughput big data systems.
- Architect and manage foundational data systems for AI infrastructure, including vector, NoSQL, and document databases.
- Build end-to-end data engineering solutions including ETL/ELT pipelines, API services, and ingestion frameworks.
- Design storage and processing layers for analytics workloads (data lakes, warehouses, distributed file systems, real-time streaming).
- Optimize and scale distributed queries and data transformations for high performance and low latency.
- Implement data quality frameworks to ensure data integrity, reliability, and governance.
What you'll need
- Extensive data engineering experience with hands-on delivery in complex data environments.
- Deep practical understanding of database ecosystems powering AI/ML infrastructure (vector, NoSQL, and document stores).
- Hands-on experience building and shipping large-scale data platforms in production.
- Deep practical experience with distributed data processing frameworks (Apache Spark, Flink, Hadoop).
- Experience with message brokers and event streaming platforms (Apache Kafka, Kinesis).
- End-to-end exposure to data pipeline lifecycle development with workflow orchestration tools (Apache Airflow, Dagster).
- Advanced SQL skills and proficiency in Python.
Details
- Open to candidates in Eastern or Central time zones.
- Work is hybrid with onsite work expected two days per week for employees within commuting distance of an office.
Read the full description and apply on the company’s own careers page.