Overview
The Senior IT Data Engineer leads the design, development and maintenance of scalable data pipelines and infrastructure for Pharma R&D data submission, content generation and reuse. The role independently builds ETL/ELT processes, optimizes data storage, ensures data quality and reliability, and collaborates with data scientists, analysts, architects, product teams and business stakeholders.
What you'll do
- Lead the design, build and maintenance of scalable data pipelines.
- Execute specific data engineering projects independently from start to finish.
- Solve complex data ingestion and processing problems.
- Optimize data flows and enhance overall system performance.
- Engage with broader business units and data scientists to understand and meet their data needs.
- Communicate with technical and non-technical stakeholders.
- Integrate data from databases, APIs, applications, files, streaming platforms and third-party systems.
- Develop data transformation and processing workflows using SQL, Python and distributed processing technologies.
- Handle big-sized data engineering projects.
- Implement strategies that improve the organization's data infrastructure.
- Manage complex, large-scale data systems.
- Integrate diverse data sources and ensure efficient, reliable data flow across platforms.
- Develop reusable engineering components.
- Follow software engineering practices including Git, CI/CD, testing, code reviews and documentation.
- Collaborate with technical leads across Data, Content, Integration, UI and Platform, as well as architects, product teams and business stakeholders, to translate business and technical requirements into scalable solutions.
What you'll need
- Demonstrated experience handling big-sized data engineering projects and managing complex, large-scale data systems.
- Proven track record of independently building ETL/ELT processes.
- Experience optimizing data storage solutions such as data warehouses and data lakes.
- Experience ensuring data quality and reliability and monitoring data systems.
- Hands-on experience with at least one major cloud platform: AWS, Azure or GCP.
- Understanding of data formats and technologies such as Parquet, JSON and Delta Lake.
- Strong understanding of data quality, reliability, security and governance principles.
- Proficiency in programming languages such as SQL, Python, Java or Scala.
- Strong, expert knowledge of cloud platforms and Big Data tools, specifically Hadoop and Spark.
- Advanced capability to solve complex data ingestion and processing problems and optimize data flows for enhanced system performance.
- Exceptional communication and collaboration skills, with the proven ability to engage effectively with technical audiences such as data scientists and non-technical business stakeholders.
Details
- Location: Hyderabad.
- Shift: CET time zone.
Read the full description and apply on the company’s own careers page.