Overview
Lead the design, development, and operation of reliable, scalable data platforms and products. Combine hands-on data engineering with technical leadership to deliver secure, well-governed data pipelines and curated datasets for analytics, reporting, and AI use cases.
What you'll do
- Lead the architecture and delivery of scalable batch and streaming data pipelines, data products, and lakehouse solutions across cloud environments.
- Translate business, reporting, analytics, and AI requirements into technical designs, data models, interfaces, and delivery plans.
- Design ingestion and transformation patterns for relational databases, APIs, event streams, files, and cloud storage, including structured and semi-structured data.
- Build and review production-grade data solutions using Python, PySpark, SQL, Spark, GCP BigQuery, Azure Databricks, and Azure Data Lake.
- Establish reusable engineering patterns for orchestration, modular transformations, metadata, schema evolution, incremental processing, and backfills.
- Guide data modeling across raw, refined, and curated layers, ensuring datasets are understandable, reusable, performant, and aligned with domain needs.
- Define and uphold standards for coding, peer review, testing, version control, CI/CD, release management, and operational readiness.
- Own reliability and performance outcomes for critical pipelines, including monitoring, alerting, recovery, capacity planning, and cost optimization.
- Implement data quality controls, reconciliation, lineage, and observability so completeness, accuracy, freshness, and consistency can be measured.
- Partner with security, governance, architecture, platform, analytics, and business teams to address access controls, privacy, retention, and compliance requirements.
- Lead migration and modernization work from legacy data platforms to cloud environments, including dependency analysis, parity validation, cutover planning, and decommissioning support.
- Mentor and coach data engineers, provide technical direction, unblock delivery, and help the team grow its engineering and platform skills.
- Break down complex initiatives into milestones, estimate effort, surface risks and dependencies early, and communicate delivery progress to stakeholders.
- Investigate and resolve complex production issues, conduct root-cause analysis, and ensure corrective actions prevent recurrence.
- Prepare governed, well-documented datasets for AI and Generative AI use cases, including retrieval, feature, and semantic search workflows where appropriate.
- Evaluate tools and design options pragmatically, document trade-offs, and recommend approaches that meet scale, security, cost, and maintainability needs.
What you'll need
- 6+ years of experience in data engineering, data platform engineering, or a closely related discipline, including ownership of production data systems.
- Demonstrated experience leading technical delivery, mentoring engineers, and coordinating work across multiple teams or domains.
- Strong programming skills in Python and advanced SQL.
- Hands-on experience with PySpark or Spark for distributed data processing.
- Deep understanding of cloud data architecture and lakehouse concepts, including storage, compute, partitioning, file formats, and workload design.
- Hands-on experience with one or more major cloud data platforms.
- Experience designing and operating reliable ETL/ELT pipelines, workflow orchestration, data transformations, and reusable ingestion frameworks.
- Strong data modeling skills, including dimensional modeling and the design of curated datasets for analytics and downstream applications.
- Experience with data quality, reconciliation, observability, lineage, metadata, and production monitoring practices.
- Working knowledge of security and governance practices such as role-based access, encryption, sensitive data handling, and auditability.
- Experience with Git-based development, automated testing, CI/CD, infrastructure or configuration management, and controlled deployments.
- Ability to troubleshoot performance and reliability issues across distributed data systems and explain technical findings clearly.
- Strong communication, planning, prioritization, and stakeholder management skills; able to make technical topics accessible to non-engineering partners.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 6+ years of relevant experience in data engineering or data platform roles, with evidence of technical leadership and successful production delivery.
- Experience owning solutions through design, implementation, deployment, and ongoing operations.
- Ability to balance hands-on technical contribution with team guidance and cross-functional collaboration.
- Willingness to learn and work with Generative AI and Large Language Models (LLMs) in enterprise data environments.
- Willingness to learn and work with preparing governed structured and unstructured data for AI applications.
- Willingness to learn and work with embeddings, vector search, semantic search, and retrieval-augmented generation (RAG).
- Willingness to learn and work with AI-assisted engineering, data discovery, and workflow automation.
- Willingness to learn and work with new cloud data services and evolving data governance and observability practices.
Nice to have
- GCP and Azure experience, including BigQuery, Databricks, and Azure Data Lake.
- An advanced degree.
Details
- Location: Bangalore, India.
Read the full description and apply on the company’s own careers page.