Overview
Build and support the analytics layer on Databricks that connects curated data to dashboards, self-service BI, and advanced analytics across a Commercial business domain. Develop well-modeled datasets, metrics, and semantic assets while working under the guidance of senior engineers.
What you'll do
- Design, build, and support analytics-ready data products, curated datasets, and reusable transformation assets within a modern lakehouse environment.
- Translate business and analytics requirements into data specifications, dimensional models, metrics definitions, and dataset readiness criteria.
- Develop and maintain scalable SQL- and Python-based transformation logic for structured and semi-structured pharma datasets, including claims, patient, sales, payer, HUB, and specialty pharmacy data.
- Apply internally developed accelerators, reusable code patterns, templates, and engineering guardrails.
- Create fact and dimension tables, cross-domain joins, slowly changing dimensions, and business-rule-driven metrics.
- Implement data quality checks, reconciliation logic, validation routines, and anomaly detection controls.
- Document data products, metric definitions, lineage, assumptions, and known limitations.
- Partner with data engineers, analysts, BI developers, and business stakeholders on dashboard and reporting use cases.
- Apply data governance standards, access controls, and compliant handling practices for sensitive and regulated data, including PII/PHI awareness.
What you'll need
- A bachelor's or master's degree in Computer Science, Engineering, Information Systems, Statistics/Mathematics, Analytics, or a related field, or equivalent practical experience.
- 1–3 years of hands-on experience in analytics engineering, data engineering, business intelligence engineering, or a related role building curated datasets and analytics-ready assets.
- Experience transforming large-scale structured and semi-structured datasets using SQL and Python, with attention to data quality and performance.
- Strong proficiency in SQL for analytics transformations, data modeling, validation, and performance tuning.
- Working knowledge of Python for data preparation, validation, automation, and analytical workflows.
- Experience developing analytics-ready datasets using dimensional modeling concepts such as facts, dimensions, grains, keys, and slowly changing dimensions.
- Familiarity with ETL/ELT patterns, lakehouse concepts, and medallion architecture, including bronze, silver, and gold/refined layers.
- Hands-on familiarity with Databricks notebooks, jobs/workflows, and Delta Lake concepts.
- Understanding of semantic layer concepts, metric definitions, and how curated datasets support BI, dashboards, and self-service analytics.
- Experience with data quality practices including profiling, cleansing, standardization, reconciliation, and anomaly detection.
- Working knowledge of governance and secure data handling practices, including documentation, lineage, and access controls.
- Familiarity with BI/visualization tools such as Tableau or Power BI.
- Understanding of Git/version control, code reviews, and basic CI/CD concepts.
- Strong problem-solving and communication skills, with the ability to work effectively with technical and business stakeholders.
- Willingness to work as part of a deployed team embedded within a business domain and adapt to domain-specific priorities.
Nice to have
- Experience with Databricks, Delta Lake, or lakehouse architecture.
- Experience working with commercial pharma datasets such as claims, sales, payer, patient, HUB, or specialty pharmacy data.
- Exposure to BI/reporting use cases, semantic layer design, or dashboard-ready data modeling.
Details
- Location: Hyderabad, Telangana, India.
- The role is part of a deployed engineering team embedded alongside an assigned Commercial business domain.
Read the full description and apply on the company’s own careers page.