Site Reliability Engineer

Posted 59mins ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Staff SRE defining reliability for Yuno’s AI-powered payments infrastructure. Owning AWS architecture, messaging, observability, incident response, and resilience at scale.

Responsibilities:

  • Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
  • Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
  • Define SLO culture, error-budget policy, and incident practices across engineering teams
  • Design and own durable, reliable asynchronous messaging for inter-service communication
  • Own cloud infrastructure and automate provisioning with Infrastructure as Code
  • Ensure the platform scales reliably as transaction volume grows
  • Build monitoring, tracing, and alerting systems for platform health
  • Serve as senior escalation point for difficult production incidents
  • Run blameless postmortems and root-cause analyses that produce permanent fixes
  • Conduct continuous fault injection and resilience experiments
  • Mentor senior and mid-level engineers and raise organization-wide reliability standards

Requirements:

  • 7+ years of experience
  • Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
  • Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
  • Experience migrating systems from synchronous to asynchronous communication
  • Deep AWS experience with EC2, VPC, IAM, S3, and RDS
  • Strong networking fundamentals
  • Infrastructure as Code experience with Terraform or Pulumi
  • Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
  • Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
  • Track record defining and operating SLOs, SLIs, and error budgets across services
  • Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
  • Distributed systems debugging experience
  • Comfortable coding automation and tooling in Go, Python, or similar
  • Solid SQL and PostgreSQL knowledge
  • NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
  • Proven technical leadership, architecture influence across teams, and engineering mentorship
  • Advanced written and spoken English proficiency

Benefits:

  • Competitive Compensation
  • Remote Work — you can work from everywhere
  • Home Office Bonus — a one-time allowance to set up your ideal home office
  • Work Equipment
  • Stock Options
  • Health Plan wherever you are
  • Flexible Days Off
  • Language, Professional, and Personal Growth courses