Overview
The Data Engineer, Specialist is responsible for building and maintaining scalable data pipelines, integrating data across platforms, and supporting analytics and business intelligence initiatives.
What you'll do
- Design, develop and maintain scalable ETL processes to extract data from structured and unstructured sources while ensuring accuracy, consistency and performance optimization.
- Architect and manage database systems for large-scale data storage and retrieval, ensuring high availability, security and efficiency with complex datasets.
- Integrate and transform data from APIs, on-premises databases and cloud storage to create unified datasets for data-driven decision-making.
- Collaborate with business intelligence analysts, data scientists and other stakeholders to understand data needs and deliver high-quality, business-relevant datasets.
- Monitor, troubleshoot and optimize data pipelines and workflows to resolve performance bottlenecks, improve processing efficiency and ensure data integrity.
- Develop automation frameworks for data ingestion, transformation and reporting to streamline data operations and reduce manual effort.
- Work with cloud-based data platforms and technologies such as AWS (Redshift, Glue, S3), Google Cloud (Big Query, Dataflow), or Azure (Synapse, Data Factory) to build scalable data solutions.
- Optimize data storage, indexing and query performance to support real-time analytics and reporting while ensuring cost-effective and high-performing data solutions.
- Lead or contribute to special projects involving data architecture improvements, migration to modern data platforms, or advanced analytics initiatives.
What you'll need
- Minimum 5 years of relevant work experience.
- At least 2 years in designing and maintaining data pipelines, ETL processes and database architectures.
- Bachelor’s degree (B.E./B.Tech) in Computer Science or IT, or bachelor’s in data science, Statistics, or related field.
- Cloud platforms experience with AWS (S3, EMR, Glue, Lambda, Kinesis, Athena, Iceberg, Step Functions, SNS/SQS, IAM, CloudWatch).
- Experience with Databricks: Delta Lake, Unity catalog volumes, Lakeflow connect, Dataframes, Apache Spark.
- Database experience with Redshift and Aurora.
- Languages and frameworks experience with Pyspark, Python, Pandas, NumPy and SQL.
- Orchestration and streaming experience with Airflow, Kafka and Lakeflow declarative pipelines.
- Monitoring and DevOps experience with Docker, Terraform, GitHub Actions, Code Pipeline, spark UI and Git integration.
- Product build skills in data design.
- Functional domain exposure from Wealth and asset management, investments or Banking domain.
- AI exposure to Claude Code or GitHub copilot.
Details
- Location: Hyderabad, Telangana, India.
- The role is based in Hyderabad, Telangana at Vanguard India.
- Vanguard has implemented a hybrid working model for most employees.
Read the full description and apply on the company’s own careers page.