Overview
We are looking for an experienced Data Engineer / Data Scientist to develop scalable data solutions, support machine learning and MLOps workflows, and collaborate with cross-functional teams while independently owning assigned deliverables.
What you'll do
- Develop scalable data processing solutions using Python, PySpark, and Azure Databricks.
- Build and maintain batch and streaming data pipelines.
- Develop and optimize Spark DataFrame-based transformations.
- Debug and troubleshoot Spark jobs in Azure Databricks.
- Implement Delta Lake solutions for data reliability, versioning, and efficient querying.
- Develop APIs using Python or Scala for data and machine learning applications.
- Support machine learning projects, MLOps workflows, and model deployment activities.
- Work with Azure services for data ingestion, storage, security, integration, and processing.
- Configure and manage Databricks job clusters and notebook workflows.
- Validate processed data by building and executing DataFrame-based checks.
- Collaborate with technical and business teams while independently owning assigned deliverables.
What you'll need
- Strong hands-on expertise in Python and PySpark.
- Good knowledge of Microsoft Azure and Azure Databricks.
- Hands-on experience with MLOps practices and tools.
- Basic understanding of machine learning model deployment.
- Practical experience working on machine learning projects.
- Experience developing and debugging Spark-based applications.
- Ability to work independently and interact effectively with project teams.
- Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, Information Technology, or a related discipline.
- 5–8 years of relevant experience in data engineering, data science, machine learning, or cloud analytics.
- Hands-on experience in Databricks notebook development.
- Strong experience with Spark DataFrames using PySpark or Scala.
- Experience debugging and optimizing Spark jobs in Azure Databricks.
- Knowledge of API development using Python or Scala.
- Working knowledge of Azure services, including Azure Event Hubs, Azure Storage Accounts, Azure Key Vault, Azure Service Bus, Azure Functions, and Azure Data Lake Storage.
- Understanding of Databricks job clusters and compute configurations.
- Hands-on experience implementing Azure cloud-based data solutions.
- Knowledge of real-time streaming technologies such as Kafka.
- Experience developing batch and streaming pipelines using Event Hubs, Kafka, or IoT data sources.
- Experience implementing Delta Lake solutions for data reliability, versioning, and query performance.
- Working knowledge of GitHub or similar version-control platforms.
Nice to have
- MLflow or similar tools for experiment tracking and model lifecycle management.
- CI/CD implementation for data and machine learning workloads.
- Performance tuning of Spark applications and Databricks workloads.
- Data quality validation, monitoring, and production support.
- Secure integration of Azure services using managed identities and secrets.
Details
- Position type: Immediate Requirement.
- Location: Bangalore, India.
- Work mode: Remote.
- Employment type: Full time.
Read the full description and apply on the company’s own careers page.