Site Reliability Engineer
Posted 59mins ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff SRE defining reliability for Yuno’s AI-powered payments infrastructure. Owning AWS architecture, messaging, observability, incident response, and resilience at scale.
Responsibilities:
- Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
- Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
- Define SLO culture, error-budget policy, and incident practices across engineering teams
- Design and own durable, reliable asynchronous messaging for inter-service communication
- Own cloud infrastructure and automate provisioning with Infrastructure as Code
- Ensure the platform scales reliably as transaction volume grows
- Build monitoring, tracing, and alerting systems for platform health
- Serve as senior escalation point for difficult production incidents
- Run blameless postmortems and root-cause analyses that produce permanent fixes
- Conduct continuous fault injection and resilience experiments
- Mentor senior and mid-level engineers and raise organization-wide reliability standards
Requirements:
- 7+ years of experience
- Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
- Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
- Experience migrating systems from synchronous to asynchronous communication
- Deep AWS experience with EC2, VPC, IAM, S3, and RDS
- Strong networking fundamentals
- Infrastructure as Code experience with Terraform or Pulumi
- Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
- Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
- Track record defining and operating SLOs, SLIs, and error budgets across services
- Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
- Distributed systems debugging experience
- Comfortable coding automation and tooling in Go, Python, or similar
- Solid SQL and PostgreSQL knowledge
- NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
- Proven technical leadership, architecture influence across teams, and engineering mentorship
- Advanced written and spoken English proficiency
Benefits:
- Competitive Compensation
- Remote Work — you can work from everywhere
- Home Office Bonus — a one-time allowance to set up your ideal home office
- Work Equipment
- Stock Options
- Health Plan wherever you are
- Flexible Days Off
- Language, Professional, and Personal Growth courses
















