Atlan logo

Senior Software Engineer - Reliability

Atlan
NewPosted today

LOCATION

India · Remote

EXPERIENCE

Not specified

TYPE

FullTime

SKILLS REQUIRED

Incident responseReliability EngineeringRoot Cause AnalysisGoFault InjectionEvaluation Harnesses

Job description

Overview

Join Atlan's Reliability team to build and improve a production multi-agent AI SRE platform. The role focuses on making agent investigations accurate and trustworthy, expanding safe auto-remediation, and enabling other teams to use the platform in production.

What you'll do

  • Improve investigation agents' root-cause accuracy so engineers can trust their answers without re-checking.
  • Expand auto-remediation from a handful of playbooks to dozens, using staged dry-run, human-approval and autonomous execution.
  • Build and run fault-injection benchmarks and evaluation harnesses for agent trustworthiness.
  • Set evaluation pass marks, report results with honest denominators, and validate the evaluation harness itself.
  • Design fail-closed checks, kill switches, approval flows and blast-radius limits for agents writing to production.
  • Build new agents that remove categories of operational toil.
  • Build clean interfaces, guardrails and safe defaults so other teams can use the reliability platform in their own production environments.
  • Respond to production issues by stopping the bleeding first and then removing the underlying class of failure.

What you'll need

  • Experience carrying a pager, owning incidents end to end, or working in a support or escalation queue long enough to understand which operational pain is worth removing.
  • Experience eliminating a recurring operational problem rather than only making a runbook faster.
  • Ability to identify who used what you built, what changed, and the relevant numbers and denominators.
  • Experience rebuilding a core part of your own work end to end with AI so that the workflow is structurally different, not just faster.
  • Experience shipping an AI-native workflow that other people depend on.
  • Ability to describe guardrails before capabilities when architecting agents that take action autonomously.
  • Ability to reason about false-positive rates, rollback paths and blast radius when agents are wrong in production.
  • Ability to build clean interfaces, document failure modes and provide safe defaults for someone else's production.
  • Ability to contribute quickly to a high-rigor existing codebase and improve it rather than rewrite it.
  • Ability to use deterministic code for routing, filtering and safety before generative steps.
  • Ability to think in systems and use agents as leverage rather than as a replacement for judgment.
  • Genuine interest in reliability work and operational efficiency.

Nice to have

  • Experience getting other teams to adopt a platform you built.

Details

  • Location: India.
  • Work best with overlap between roughly 11am and 8pm IST.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Senior Software Engineer - Reliability