Overview
Build and operate an AI/ML platform for root cause analysis (RCA) that uses vector search, RAG retrieval, and confidence scoring to decide when to notify engineers automatically versus escalating to humans.
What you'll do
- Design and operate a pgvector schema for tenant-isolated incident and operational context data.
- Build and maintain embedding pipelines to convert operational documents into searchable vectors with chosen chunking strategies.
- Implement a RAG retrieval layer that tunes relevance thresholds, re-ranking, and hybrid dense+sparse search for incident-time document selection.
- Design the ingestion lifecycle including runbook imports, historical backfill, CI/CD post-deploy hooks, and feedback loops for resolved incidents.
- Assemble and iterate a prompt that combines Grafana stateless RCA inputs, live signal data, and retrieved context for structured RCA output.
- Define and enforce an RCA JSON schema using Pydantic validation and route results via FastAPI gate logic.
- Own confidence scoring calibration (0–1) and evaluation via golden test sets, judge-prompt scoring, and monthly calibration (ECE).
What you'll need
- 4+ years of experience in AI/ML engineering, data engineering, NLP, or a closely related field.
- Hands-on LLM work.
- Production experience building and operating a vector database (pgvector, Pinecone, Weaviate, or equivalent).
- Strong understanding of embedding models and how embedding quality affects retrieval accuracy.
- Experience designing and operating data ingestion pipelines at scale (Celery, Airflow, or similar).
- Demonstrated production prompt engineering for reliable structured outputs.
- Strong Python skills including FastAPI, Pydantic, SQLAlchemy, and async patterns.
Details
- Experience with multi-tenant architectures and strict tenant isolation is required.
Read the full description and apply on the company’s own careers page.