Overview
Design, build, and support scalable data pipelines for business intelligence and analytics platforms supporting trading operations. Work with Trading Analytics & Insights in an Agile delivery environment using Kanban, Scrum, and Azure DevOps.
What you'll do
- Architect, design, and implement robust data pipelines using Python, Pandas, SQL, Apache Airflow, and Databricks.
- Build and maintain ETL/ELT processes that ingest data from APIs, relational databases, event streams, and flat files.
- Develop cloud-based data solutions using AWS services such as S3, EC2, EKS, IAM, Glue, and CloudWatch where appropriate.
- Optimise, monitor, and troubleshoot pipelines for performance, cost, resilience, and scalability.
- Use distributed processing technologies such as Spark/PySpark and Delta Lake for large-scale workloads.
- Implement data quality checks, lineage, observability, and clear operational runbooks.
- Create clear data visualisations and technical communication for technical and non-technical audiences.
- Work with multi-functional partners to translate trading and business needs into fit-for-purpose data solutions.
- Apply Agile practices, version control, code review, automated testing, and release subject area.
- Use approved AI tools to accelerate delivery while protecting company data, validating output, and retaining accountability for production decisions.
- Use coding assistants and AI tools for code drafts, tests, documentation, SQL review, debugging, and knowledge discovery.
- Write clear prompts stating the goal, constraints, inputs, expected output, and acceptance checks.
- Create useful context for AI tools using relevant schemas, examples, business rules, and bounded source material without exposing critical information.
- Evaluate AI-generated work through testing, peer review, source checks, and security review before production use.
- Keep current with approved emerging AI capabilities and propose practical uses that improve productivity or data quality.
What you'll need
- 3+ years of data engineering experience in an enterprise environment.
- Strong Python and Pandas skills, including HTTP/API integration, for example Requests.
- Use web automation only where it is approved, lawful, and appropriate for the source.
- Advanced SQL and solid understanding of relational data modelling and databases.
- Proven ETL/ELT and pipeline design experience, including hands-on Apache Airflow.
- Hands-on Databricks and distributed processing experience.
- Experience building and operating data solutions on AWS, including S3 and relevant managed data services.
- Familiarity with IAM.
- Solid understanding of Linux, Git, pull-request workflows, and CI/CD.
- Familiarity with Docker, infrastructure as code, for example Terraform and Jenkins, and secure cloud deployment practices.
- Experience with Tableau, Power BI, or similar visualisation tools.
- Understanding of secure engineering standards and documentation practices.
- Basic practical knowledge of LLMs, RAG, prompt design, context design, and safe use of AI developer tools.
- Strong analytical, problem-solving, communication, collaboration, and organisational skills.
- A responsible, self-motivated, adaptable approach and a genuine interest in automation.
- A Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
Nice to have
- Spark/PySpark and Delta Lake are strongly preferred.
- Experience with EKS is helpful where workloads run on Kubernetes.
- Experience with messaging or streaming systems such as Kafka; familiarity with event-driven data patterns is helpful.
- Azure DevOps experience is helpful.
- Experience in trading, financial services, energy, or another regulated data-intensive domain is advantageous.
Details
- Location: Pune, India.
- This position is a hybrid of office/remote working.
- Up to 10% travel should be expected with this role.
- This role is eligible for relocation within country.
Read the full description and apply on the company’s own careers page.