Overview
GCP Data Engineer responsible for architecting, developing, and optimizing scalable GCP-based data pipelines and data management solutions.
What you'll do
- Design, develop, and maintain scalable, resilient data pipelines on GCP for analytics and reporting.
- Collaborate with business stakeholders, data scientists, and analytics teams to implement data workflows based on requirements.
- Implement and enforce data quality, security, and governance standards.
- Optimize data ingestion, transformation, and processing for high performance and cost efficiency.
- Perform data profiling, troubleshooting, and resolution of pipeline issues to ensure operational reliability.
- Document architecture, workflows, procedures, and operational guidelines.
What you'll need
- Minimum 6 years of practical experience in data engineering with a significant focus on GCP and Big Data ecosystems.
- Extensive experience with GCP services including BigQuery, Dataflow, Cloud Storage, and Cloud Pub/Sub.
- Strong proficiency in Apache Spark for distributed data processing and analytics.
- Hands-on experience building and maintaining data pipelines using ETL/ELT processes.
- Proficiency in Python for data scripting, automation, and orchestration.
- Experience with distributed data storage and management including PostgreSQL, MySQL, and MongoDB.
- Working knowledge of Git and Linux/Unix environments for data processing and scripting.
- Familiarity with version control tools such as Git.
- Experience with data governance, metadata management, and data security best practices.
- Experience with data orchestration tools like Apache Airflow or Prefect.
- Understanding of containerization (Docker) and orchestration (Kubernetes) for data deployment.
Details
Read the full description and apply on the company’s own careers page.