Senior Software Engineer – Reliability
Posted 3hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Software Engineer improving Atlan's multi-agent AI SRE platform for enterprise data and AI context. Building trusted investigations, safe auto-remediation, evaluation harnesses, and production guardrails.
Responsibilities:
- Improve investigation-agent root-cause accuracy so engineers can trust diagnoses without re-checking
- Expand auto-remediation from a handful of playbooks to dozens using staged autonomy
- Build and run fault-injection benchmarks and evaluation harnesses
- Set evaluation pass marks, report results with honest denominators, and validate the harnesses
- Design fail-closed checks, kill switches, approval flows, and blast-radius limits for production actions
- Build new agents that remove categories of operational toil
- Develop clean interfaces, guardrails, and safe defaults for other teams using the reliability platform
- Stop production incidents from worsening and permanently remove recurring failure classes
- Contribute to Atlan's mature, opinionated codebase and improve its documented architecture and eval-gated pull requests
- Collaborate across teams to drive adoption of the reliability platform
Requirements:
- Experience carrying a pager, owning incidents end to end, or working a support or escalation queue
- Demonstrated ability to eliminate recurring operational problems rather than merely optimize runbooks
- Experience measuring outcomes using adoption metrics, real numbers, and honest denominators
- Experience rebuilding a core part of one's work with AI and shipping an AI-native workflow used by others
- Experience architecting agents that take autonomous action, including defining guardrails
- Understanding of false-positive rates, rollback paths, and blast radius
- Experience building platforms for other teams, including clean interfaces, documented failure modes, and safe defaults
- Ability to contribute quickly to a high-rigor existing codebase and improve it without rewriting it
- Ability to write deterministic code for routing, filtering, and safety before generative steps
- Genuine interest in reliability and operational efficiency
- Availability for overlap between roughly 11am and 8pm IST
Benefits:
- Strong base salary
- Performance-based variable pay
- Impact-driven equity for most roles
- Health, dental, vision, and mental health benefits from Day 1
- Flexible health stipends
- Flexible time off
- Modern leave policies
- Accelerated growth and learning opportunities
- Global, remote-first, high-trust work environment
- Work from anywhere with a diverse team across 15+ countries
- Async work environment
- Flexibility and ownership over how you work


















