Citigroup logo

PySpark Data Engineer – Assistant Vice President

Citigroup
NewPosted yesterday

LOCATION

Chennai · Hybrid

EXPERIENCE

9 - 11 Years

TYPE

FullTime

SKILLS REQUIRED

PySpark Data EngineeringSQLData ModelingData PipelinesETL DevelopmentData GovernanceCI/CD

Job description

Overview

Develop and support high-quality, scalable data products and analytical data models for regulatory requirements and data-driven decision making. Serve as an example to team members, work with customers, remove or escalate roadblocks, and contribute to business outcomes on an agile team.

What you'll do

  • Develop and support scalable, extensible, and highly available data solutions.
  • Deliver on critical business priorities while aligning with the wider architectural vision.
  • Identify and help address potential risks in the data supply chain.
  • Follow and contribute to technical standards.
  • Design and develop analytical data models.

What you'll need

  • First Class Degree in Engineering/Technology (4-year graduate course).
  • 9 to 11 years’ experience implementing data-intensive solutions using agile methodologies.
  • Experience of relational databases and using SQL for data querying, transformation and manipulation.
  • Experience of modelling data for analytical consumers.
  • Ability to automate and streamline the build, test and deployment of data pipelines.
  • Experience in cloud native technologies and patterns.
  • A passion for learning new technologies and a desire for personal growth through self-study, formal classes, or on-the-job training.
  • Excellent communication and problem-solving skills.
  • An inclination to mentor and an ability to lead and deliver medium sized components independently.
  • Hands-on experience of building ETL data pipelines.
  • Proficiency in two or more data integration platforms such as Ab Initio, Apache Spark, Talend and Informatica.
  • Experience of big data platforms such as Hadoop, Hive or Snowflake for data storage and processing.
  • Expertise around data warehousing concepts and relational database design using Oracle, MSSQL or MySQL, and NoSQL database design using MongoDB or DynamoDB.
  • Good exposure to data modeling techniques, including design, optimization and maintenance of data models and data structures.
  • Proficiency in one or more programming languages commonly used in data engineering such as Python, Java or Scala.
  • Exposure to DevOps concepts and enablers, including CI/CD platforms, version control and automated quality control management.
  • A strong grasp of data governance principles and practice, including data quality, security, privacy and compliance.

Nice to have

  • Experience developing Ab Initio Co>Op graphs and ability to tune for performance.
  • Demonstrable knowledge across the full suite of Ab Initio toolsets, including GDE, Express>IT, Data Profiler, Conduct>IT, Control>Center and Continuous>Flows.
  • Good exposure to public cloud data platforms such as S3, Snowflake, Redshift, Databricks and BigQuery, and demonstratable understanding of underlying architectures and trade-offs.
  • Exposure to data validation, cleansing, enrichment and data controls.
  • Fair understanding of containerization platforms like Docker and Kubernetes.
  • Exposure to working on event, file or table formats such as Avro, Parquet, Protobuf, Iceberg and Delta.
  • Experience using a job scheduler such as Autosys.
  • Exposure to business intelligence tools such as Tableau and Power BI.
  • Certification on any one or more of the above topics would be an advantage.

Details

  • Location: Chennai, Tamil Nadu, India.
  • Employment type: Full time.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

PySpark Data Engineer – Assistant Vice President