Senior Software Engineer – Reliability

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior Software Engineer improving Atlan's multi-agent AI SRE platform for enterprise data and AI context. Building trusted investigations, safe auto-remediation, evaluation harnesses, and production guardrails.

Responsibilities:

  • Improve investigation-agent root-cause accuracy so engineers can trust diagnoses without re-checking
  • Expand auto-remediation from a handful of playbooks to dozens using staged autonomy
  • Build and run fault-injection benchmarks and evaluation harnesses
  • Set evaluation pass marks, report results with honest denominators, and validate the harnesses
  • Design fail-closed checks, kill switches, approval flows, and blast-radius limits for production actions
  • Build new agents that remove categories of operational toil
  • Develop clean interfaces, guardrails, and safe defaults for other teams using the reliability platform
  • Stop production incidents from worsening and permanently remove recurring failure classes
  • Contribute to Atlan's mature, opinionated codebase and improve its documented architecture and eval-gated pull requests
  • Collaborate across teams to drive adoption of the reliability platform

Requirements:

  • Experience carrying a pager, owning incidents end to end, or working a support or escalation queue
  • Demonstrated ability to eliminate recurring operational problems rather than merely optimize runbooks
  • Experience measuring outcomes using adoption metrics, real numbers, and honest denominators
  • Experience rebuilding a core part of one's work with AI and shipping an AI-native workflow used by others
  • Experience architecting agents that take autonomous action, including defining guardrails
  • Understanding of false-positive rates, rollback paths, and blast radius
  • Experience building platforms for other teams, including clean interfaces, documented failure modes, and safe defaults
  • Ability to contribute quickly to a high-rigor existing codebase and improve it without rewriting it
  • Ability to write deterministic code for routing, filtering, and safety before generative steps
  • Genuine interest in reliability and operational efficiency
  • Availability for overlap between roughly 11am and 8pm IST

Benefits:

  • Strong base salary
  • Performance-based variable pay
  • Impact-driven equity for most roles
  • Health, dental, vision, and mental health benefits from Day 1
  • Flexible health stipends
  • Flexible time off
  • Modern leave policies
  • Accelerated growth and learning opportunities
  • Global, remote-first, high-trust work environment
  • Work from anywhere with a diverse team across 15+ countries
  • Async work environment
  • Flexibility and ownership over how you work