Overview
Join Atlan's Reliability team to build and improve a production multi-agent AI SRE platform. The role focuses on making agent investigations accurate and trustworthy, expanding safe auto-remediation, and enabling other teams to use the platform in production.
What you'll do
- Improve investigation agents' root-cause accuracy so engineers can trust their answers without re-checking.
- Expand auto-remediation from a handful of playbooks to dozens, using staged dry-run, human-approval and autonomous execution.
- Build and run fault-injection benchmarks and evaluation harnesses for agent trustworthiness.
- Set evaluation pass marks, report results with honest denominators, and validate the evaluation harness itself.
- Design fail-closed checks, kill switches, approval flows and blast-radius limits for agents writing to production.
- Build new agents that remove categories of operational toil.
- Build clean interfaces, guardrails and safe defaults so other teams can use the reliability platform in their own production environments.
- Respond to production issues by stopping the bleeding first and then removing the underlying class of failure.
What you'll need
- Experience carrying a pager, owning incidents end to end, or working in a support or escalation queue long enough to understand which operational pain is worth removing.
- Experience eliminating a recurring operational problem rather than only making a runbook faster.
- Ability to identify who used what you built, what changed, and the relevant numbers and denominators.
- Experience rebuilding a core part of your own work end to end with AI so that the workflow is structurally different, not just faster.
- Experience shipping an AI-native workflow that other people depend on.
- Ability to describe guardrails before capabilities when architecting agents that take action autonomously.
- Ability to reason about false-positive rates, rollback paths and blast radius when agents are wrong in production.
- Ability to build clean interfaces, document failure modes and provide safe defaults for someone else's production.
- Ability to contribute quickly to a high-rigor existing codebase and improve it rather than rewrite it.
- Ability to use deterministic code for routing, filtering and safety before generative steps.
- Ability to think in systems and use agents as leverage rather than as a replacement for judgment.
- Genuine interest in reliability work and operational efficiency.
Nice to have
- Experience getting other teams to adopt a platform you built.
Details
- Location: India.
- Work best with overlap between roughly 11am and 8pm IST.
Read the full description and apply on the company’s own careers page.