Overview
Lead the design and technical direction of TaskUs’ modern data ecosystem, ensuring it is high-performance, cost-effective, and secure as business needs evolve.
What you'll do
- Lead architectural decisions for compute engine selection, open-table formats, and tiered storage design.
- Architect a decoupled data environment to support interoperability and reduce proprietary vendor lock-in.
- Design and oversee automated data governance, including PII discovery, row/column-level security, and auditability.
- Define pipeline “Definition of Done,” including coding standards, CI/CD patterns, and documentation requirements.
- Maintain version-controlled Architectural Decision Records (ADRs) documenting rationale and trade-offs.
- Monitor and optimize platform performance and spend to achieve sub-second query speeds while keeping a lean cloud footprint.
- Conduct deep-dive code/design reviews and mentor senior engineering staff on data models and orchestration workflows.
What you'll need
- 8+ years of data engineering or architecture experience.
- Production-grade lakehouse delivery experience for high-concurrency organizations (1,000+ users).
- Experience with Databricks (Lakehouse/Unity Catalog) and warehouses such as Amazon Redshift or Snowflake.
- Hands-on expertise with Apache Iceberg or Delta Lake, including optimization for partitioning and schema evolution.
- Mastery of dbt (Core) plus PySpark or Python for data transformation and processing.
- Advanced familiarity with Apache Iceberg or equivalent open-table format technologies (Apache Iceberg stated; open-table formats with Delta Lake stated).
- Advanced experience with Apache Airflow for resilient, dependency-aware DAGs.
Details
Read the full description and apply on the company’s own careers page.