Overview
Data Engineer II responsible for building and operating enterprise data ingestion, lakehouse modeling, data quality and observability capabilities, production support, automation, and AI systems engineering components.
What you'll do
- Build and operate ingestion pipelines from cloud and on-premise ERP, CRM, ITSM, billing/revenue, planning and quoting platforms, and third-party operational systems into the enterprise data lakehouse.
- Select and implement CDC, orchestrated batch, event streaming, packaged analytics/replication tooling, and API-based ingestion patterns based on latency, volume, and cost profiles.
- Own new-source onboarding end to end, including source analysis, schema and field availability confirmation, ingestion design, historical backfill, incremental load strategy, and exposure through curated serving layers.
- Build and extend raw landing, cleansed/conformed, and curated business layers in the medallion architecture.
- Build conformed star-schema facts and dimensions and denormalized, query-optimized serving layers for BI tools and downstream applications.
- Write and optimize SQL/stored-procedure and Spark-based transformations.
- Model dimensional entities and conformed keys for customer, supplier, product, order, contract, and site data.
- Build the semantic and reporting layer with reusable business definitions, hierarchies, metric logic, and documented ownership.
- Build and extend source-to-target reconciliation frameworks and content-level synchronization validation.
- Implement data accuracy, completeness, and freshness/recency metrics and surface them through the platform's observability layer.
- Investigate and resolve production data disconnects, including root-cause isolation, repair or reprocessing, and corrective controls.
- Tune long-running transformations, reduce compute consumption, and contribute to platform right-sizing, storage cleanup, and cost optimization.
- Manage schema drift and upstream release changes, assess downstream impact, and regression-test pipelines across lower environments before production.
- Work within CI/CD and Git-based change control with automated test generation and validation.
- Support production monitoring, alerting, incident triage, and root-cause analysis for owned pipelines, including month-end and quarter-close critical windows.
- Automate period-close processes, report generation, reconciliation, environment and admin workflows, infrastructure-as-code, and test generation.
- Assemble facts, history, retrieved knowledge, and state for each model call while favoring signal over volume.
- Build runtime components around models, including tool/action interfaces, state handling, retries, error handling, and stop conditions.
- Write automated evaluations, groundedness checks, and regression tests.
- Apply input/output constraints, permission gates, and human-in-the-loop review per platform standards.
- Instrument tracing and telemetry, investigate failures, and understand token and compute spend.
Details
- Location: Gurgaon, India.
Read the full description and apply on the company’s own careers page.