Overview
Design and operate data infrastructure for AI systems, including retrieval-augmented generation, vector databases, semantic layers, and production data pipelines.
What you'll do
- Design and scale Retrieval-Augmented Generation pipelines.
- Manage vector databases such as Pinecone, Milvus, or Weaviate.
- Build knowledge graphs and semantic layers.
- Automate data guardrails for noise, bias, and PII detection.
- Resolve data quality issues at their source.
- Build, deploy, and maintain CI/CD pipelines for data infrastructure.
What you'll need
- Expertise in data mining, data storage, and ETL processes.
- Experience developing data pipelines with tools such as Glue, Databricks, Synapse, or Dataproc.
- Experience with relational and NoSQL databases, including PostgreSQL, DB2, and MongoDB.
- Strong problem-solving, analytical, and critical-thinking skills.
- Ability to manage multiple projects with attention to detail.
- Ability to translate business needs into technical requirements.
Nice to have
- Experience as a Data Engineer or in cloud modernization.
- Data modelling experience.
- Professional or cloud platform certification.
- Familiarity with GitHub and Visual Studio.
- Degree in Computer Science, Software Engineering, Information Technology, or another scientific discipline.
Details
- Location: Bangalore, Karnataka, India.
- Hybrid-friendly culture.
Read the full description and apply on the company’s own careers page.