Overview
Build, operate, and support production data pipelines and cloud data platforms, translating complex business and financial logic into correct, testable, and maintainable code. The role focuses on AWS data engineering, data reconciliation, operational reliability, and cloud lakehouse migration.
What you'll do
- Build and operate production data pipelines on AWS using Glue (PySpark), Step Functions, Lambda, Athena, SNS, and S3, from ingestion through transformation to reconciled, reportable output.
- Own operational support for production data pipelines, including monitoring, incident management, PagerDuty response, root cause analysis, and timely resolution of production issues.
- Continuously improve platform reliability and observability.
- Develop and refactor transformation logic in PySpark, including migrating legacy Pandas/SQL logic and performance-tuning large-scale Spark jobs.
- Design data reconciliation checks, including parity between old and new implementations; investigate discrepancies to root cause; and produce evidence for business sign-off.
- Translate complex business and financial logic, including asset classification, valuation and pricing rules, and methodology calculations, into correct, testable, well-documented code.
- Take features through production readiness, including pre-checks, access and role setup, environment promotion, and clear runbooks for ongoing support.
- Contribute to the cloud lakehouse migration by modelling refined tables, building repeatable ETL, and lifting reporting workloads onto a modern Delta/Databricks platform.
- Monitor scheduled jobs, triage failures involving partitioning, schema/data-type issues, and backfills, and keep periodic month-end and quarter-end processing on time.
What you'll need
- 5+ years in data engineering with strong Python and SQL.
- Hands-on PySpark / Spark experience for large-scale data transformation, including debugging and performance tuning.
- Production experience with AWS data services: Glue, Athena, S3, Lambda, and Step Functions, or close equivalents on another major cloud.
- Solid data modelling skills and a rigorous, reconciliation-first approach to data quality.
- Ability to read a business or technical specification and turn it into correct, maintainable pipeline logic.
- Comfort with Git-based workflows, CI/CD, and infrastructure-as-code.
Nice to have
- Databricks / lakehouse (Delta Lake) experience.
- Background in financial services, including custody, fund accounting, unit pricing, or regulatory/financial reporting.
- Experience with access governance and IAM, such as role provisioning, and writing operational runbooks.
- Track record migrating legacy reporting or ETL onto modern cloud data platforms.
Details
- Based in Hyderabad, Telangana, India.
- Vanguard has implemented a hybrid working model for most employees.
Read the full description and apply on the company’s own careers page.