Overview
Senior Data Architect for AWS & Databricks modernization, leading hands-on lakehouse architecture and large-scale migrations for Clinical and Non-Clinical data.
What you'll do
- Own technical roadmap and execution for migrating legacy warehouses, on-prem databases, and point platforms to Databricks lakehouse.
- Build migration pipelines and re-platforming accelerators (schema conversion, historical backfill, dual-run validation, cutover automation).
- Define reusable modernization patterns including landing zone design and medallion conventions.
- Architect and build the Clinical/Non-Clinical lakehouse on Databricks using Delta Lake and medallion design.
- Build and optimize Delta Live Tables pipelines, Databricks Workflows, and PySpark/Spark SQL jobs.
- Own Unity Catalog governance design (catalogs, schemas, access control, lineage, and data sharing) and data quality for clinical data.
What you'll need
- 5+ years hands-on experience migrating or modernizing legacy data warehouses/on-prem platforms to a Databricks lakehouse.
- 5+ years hands-on building on Databricks (Delta Lake, Delta Live Tables, Unity Catalog, Databricks Workflows, Databricks SQL, MLflow) with production pipelines.
- Hands-on schema conversion, historical backfill, dual-run/parallel validation, and cutover automation for large-scale migrations.
- Strong PySpark and Spark SQL skills, including performance tuning and cost optimization at scale.
- Experience standardizing workspace/catalog/environment topology across dev/test/prod and multiple workspaces.
- Proven experience defining and delivering a semantic modeling/analytics-ready data roadmap deployed in practice.
- Demonstrated current use of AI-assisted/agentic tools for migration analysis, pipeline design, and documentation.
Details
- Location: Bengaluru, India.
- Role described as split 70% hands-on technical execution and 30% strategy.
Read the full description and apply on the company’s own careers page.