Overview
Lead Data Engineer responsible for designing, delivering, and owning enterprise-scale data products and pipelines for the ERP Enterprise Data Product team. The role supports SAP S/4HANA data replication into Google BigQuery and integration with the Enterprise Lakehouse on Databricks for AI/ML, conversational analytics, and global reporting.
What you'll do
- Independently design and implement complex end-to-end data pipelines spanning multiple sources and large SAP datasets, including Configuration, Master Data, and Transactional data.
- Anticipate scaling needs, data skew, and growth using partitioning, clustering, and incremental backfills.
- Act as the primary gatekeeper for technical standards.
- Ensure junior engineers are proficient in foundational tooling such as Git branching and environment setup.
- Ensure testing patterns are consistently applied.
- Own routine production incidents in familiar pipelines, including schema drift, late files, and logic errors.
- Deliver permanent fixes by updating code, tests, and documentation.
- Work directly with analysts and data scientists to refine requirements into concrete technical stories.
- Proactively surface edge cases such as late-arriving data and schema changes.
- Ensure adoption of Elanco’s data standards and foundational tools, including GitHub, VS Code, SonarQube, and Copilot, across the team.
- Use code reviews as a teaching tool.
- Monitor key pipelines using logs and dashboards.
- Identify and fix performance or memory issues using query plans and Spark UI without assistance.
- Design and implement meaningful data quality checks, including referential integrity and distribution checks.
- Ensure tests run in CI/CD schedules.
- Proactively identify and fix knowledge silos by creating reusable assets and documentation.
- Contribute to the Data Engineering community across Elanco.
- Demonstrate a growth and sharing mindset, leveraging AI to accelerate intermediate tasks and tackle complex engineering challenges.
What you'll need
- Bachelor’s degree in computer science, Software Engineering, or equivalent professional experience.
- 6+ years of experience engineering and delivering enterprise scale data solutions.
- Proven experience with Google BigQuery, including BigQuery Studio.
- Proven experience with Databricks, including PySpark and Delta Lake.
- Experience mentoring junior engineers, stewarding a practitioner community, and/or ensuring a team’s adherence to technical standards.
Nice to have
- Familiarity with SAP data structures, including S/4HANA.
- Familiarity with SAP BW.
- Experience with replication tools such as SAP SLT (System Landscape Transformation Server).
- Proven ability to independently deliver moderately complex data projects within a defined design.
- Strong proficiency in SQL, including multi-join and window functions.
- Strong proficiency in Python.
- Strong proficiency in Spark for building and maintaining data products.
- Experience working with modern data architectures, including Lakehouse Bronze/Silver/Gold, star schemas, and data contracts.
- Experience working within a DevSecOps culture, including Git, CI/CD, and Test-Driven Development (TDD).
- Familiarity with data quality frameworks, data governance, and cost/efficiency considerations in cloud environments.
- Experience working in diverse global landscapes, including business, technology, regulatory, and geography.
- Excellent interpersonal and communication skills.
- Proven ability to turn rough requirements into clear, testable specifications.
Details
- Location: Bangalore, India.
- Hybrid work environment.
- Travel: 0%.
Read the full description and apply on the company’s own careers page.