Overview
Curate and govern the context layer for retrieval-augmented generation (RAG) to improve answer quality and reduce hallucinations, while protecting data/PII.
What you'll do
- Curate and label content from enterprise sources using APIs and automation.
- Define chunking and metadata schemas, labeling guidelines, and golden Q&A/evaluation sets.
- Implement chunking strategies for code repositories, documentation, tickets, and test cases.
- Run A/B experiments across vector stores and monitor answer quality vs cost/latency.
- Analyze failure cases and propose data-driven improvements.
- Enforce data governance including minimization, retention, access controls, and lineage/approvals per RAI.
What you'll need
- 3+ years of data/ML experience with embeddings and retrieval expertise.
- Strong documentation and runbook skills.
- Experience with content transformation, metadata extraction, and labeling workflows.
- Familiarity with privacy and data governance principles.
- Hands-on experience with vector stores (OpenSearch, pgvector, Kendra, or Chroma).
- Experience with REST APIs and data extraction from enterprise systems.
- Python proficiency for data pipelines and automation.
Nice to have
- Experience designing golden datasets and evaluation pipelines.
- AWS Bedrock Knowledge Bases experience.
- Familiarity with software development lifecycle and technical documentation patterns.
Details
- Location: Bengaluru, India.
- Employment type: Full time regular employee.
Read the full description and apply on the company’s own careers page.