Overview
Data Engineer - ETL builds and maintains data pipelines, data warehouses, and data lakes to keep data accurate, accessible, and secure.
What you'll do
- Build and maintain data architecture pipelines to transfer and process consistent data.
- Design and implement data warehouses and data lakes with required security and appropriate data volumes and velocity.
- Develop processing and analysis algorithms for the intended data complexity and volumes.
- Collaborate with data scientists to build and deploy machine learning models.
- Design, develop, and deliver scalable data solutions and robust data pipelines for large-scale transformations.
- Enable data processing and support enterprise ETL/ELT initiatives across the organisation.
What you'll need
- Experience designing, developing, and maintaining scalable data pipelines using PySpark.
- Experience building and optimising ETL/ELT solutions for large-scale enterprise data processing.
- Experience developing reusable frameworks, components, and standards for data onboarding and analytics delivery.
- Experience implementing data quality controls, validation frameworks, and reconciliation processes.
- Ability to deliver high-quality code with engineering best practices, coding standards, and automated testing.
- Strong programming skills in Python/PySpark.
Nice to have
- Hands-on experience with Databricks and/or Snowflake.
- Strong experience with AWS Cloud services including S3, Glue, EMR, Lambda, EC2, DynamoDB, IAM, CloudWatch, and CloudTrail.
- Hands-on experience optimising AWS Glue ETL jobs using PySpark.
- Experience with large-scale distributed data processing.
- Solid understanding of Data Lake, Lakehouse, and Modern Data Platform architectures.
- Expertise with Git-based source control platforms.
- Experience with Airflow.
Details
- Role based out of Bengaluru.
Read the full description and apply on the company’s own careers page.