Overview
Data Engineer to design, develop, and optimize large-scale data processing applications using Scala, Apache Spark, and Java in a hybrid setup.
What you'll do
- Design, develop, and optimize large-scale data processing applications using Scala, Apache Spark, and Java.
- Build and maintain high-performance data pipelines for processing 80–90 million records daily.
- Develop and integrate Native APIs and data services for business-critical applications and analytics.
- Optimize Spark jobs, data workflows, and distributed computing for high throughput and low latency.
- Implement best practices for data quality, monitoring, governance, and operational excellence.
- Troubleshoot and resolve performance bottlenecks across data processing and ingestion pipelines.
- Participate in code reviews and contribute to engineering standards and architecture decisions.
What you'll need
- Strong hands-on experience in Scala, Apache Spark, and Java.
- Proven experience building and supporting large-scale distributed data processing systems for tens of millions of records daily.
- Expertise in developing and consuming Native APIs and microservices.
- Strong understanding of Spark architecture and performance tuning, including partitioning and caching.
- Experience with data modeling, ETL/ELT, and large-scale batch and streaming data pipelines.
- Solid understanding of distributed systems, concurrency, and high-volume data processing.
- Experience with SQL and relational/non-relational databases.
Nice to have
- Experience in enterprise-scale data environments in Financial Services, Payments, or FinTech domains.
- Exposure to cloud platforms (AWS, Azure, or GCP) and containerized deployments.
- Familiarity with Kafka, Airflow, Hadoop ecosystem, or similar big data technologies.
- Experience with CI/CD pipelines, DevOps practices, and Agile methodologies.
Details
- Position: Data Engineer.
- Experience: 4+ years.
- Work mode: Hybrid (Capco Office).
Read the full description and apply on the company’s own careers page.