Overview
Principal Data Engineer for the ML/AI platform, leading the design, development, and governance of scalable GCP data solutions for machine learning, AI, and GenAI use cases.
What you'll do
- Design end-to-end ML-ready data architectures on GCP using services like BigQuery, GCS, Dataproc, Composer, Dataform, and Data Fusion.
- Build and standardize batch and near-real-time data pipelines for model training, feature engineering, inference, and analytics reporting.
- Define reusable patterns for ingestion, transformation, quality validation, metadata capture, lineage, and observability.
- Design data models for BI/reporting and for ML/AI/GenAI workloads across structured, semi-structured, and unstructured data.
- Partner with Data Science and ML Engineering to deliver trusted, discoverable, and reproducible datasets for experimentation and production.
- Provide technical leadership by mentoring reviewers, setting standards with GitHub/CI/CD, and escalating complex pipeline issues.
- Define and monitor SLAs/SLOs, and ensure governance, reliability, scalability, and observability in production data environments.
What you'll need
- 10+ years of experience in data engineering, data platforms, or data warehousing, including cloud-native architectures.
- Proven ability to design and implement large-scale data pipelines and platforms on GCP.
- Experience supporting production ML/AI solutions with Data Scientists and ML Engineers.
- Hands-on delivery of pipelines for model training, feature engineering, inference, or ML monitoring workflows.
- Strong skills in Python, SQL, Spark, and distributed data processing frameworks.
- Deep experience with GCP services including BigQuery, GCS, Dataproc, Composer, Dataform, and Data Fusion.
- Strong understanding of data architecture patterns such as ELT, CDC, data products, lakehouse concepts, and domain-oriented architectures.
Details
- Work location: Bangalore, India.
Read the full description and apply on the company’s own careers page.