Site Reliability Engineer

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer operating Sporttrade’s regulated sports betting exchange across cloud and datacenter infrastructure. Automating operations, improving observability, and leading incident response for a live marketplace.

Responsibilities:

  • Own the daily operation of the exchange trading lifecycle, including market startup and shutdown, enabling and disabling trading, pre- and post-session sanity checks, and capture of settlement, clearing, and trade-reporting artifacts
  • Participate in an on-call rotation for a live regulated marketplace
  • Lead incident response, drive incidents to resolution, author postmortems, and convert one-off fixes into runbooks and automation
  • Operate and improve the observability stack, including dashboards, alert quality, SLOs, and time-to-detection
  • Run and maintain hybrid infrastructure across Kubernetes clusters, cloud accounts, and geographically distributed on-premises datacenters
  • Automate infrastructure and operational procedures using Ansible, Terraform, and Jenkins pipelines, with secrets managed in HashiCorp Vault
  • Support the exchange data platform, including PostgreSQL, Kafka change-data-capture and streaming pipelines, Redis, backup/restore, and disaster recovery
  • Support market-maker and partner connectivity, conformance testing, and partner onboarding
  • Contribute to process improvement and establish policies and procedures for monitoring, incident management, change control, and exchange operations

Requirements:

  • 5+ years of experience in a Site Reliability Engineering, DevOps, production engineering, or technical operations role supporting a 24/7 production system
  • Strong Linux fundamentals and scripting ability
  • Experience supporting and debugging Java applications in production, including stack traces, thread dumps, JVM memory, garbage collection, logs, and metrics
  • Solid working knowledge of TCP/IP networking, including connections, ports, routing, and firewalls
  • Hands-on experience operating Kubernetes in production
  • Experience managing infrastructure as code with Terraform and Ansible
  • Experience with CI/CD pipelines using Jenkins or similar
  • Experience with observability tooling such as Datadog, Prometheus, and Grafana, or equivalents
  • Track record of being on-call for critical systems
  • Working knowledge of SQL and relational databases; PostgreSQL preferred
  • Self-starter able to deliver results with minimal guidance
  • Comfortable working independently and with a team
  • Excellent communication and organizational skills, especially written incident communication and documentation
  • Background or interest in trading, capital markets, exchange operations, or sports betting is a plus
  • Familiarity with exchange protocols is a plus
  • Previous experience in a regulated industry is a plus
  • Startup experience preferred but not required

Benefits:

  • Medical, Dental, and Vision Benefits: Company pays 100% Employee premium and 50% Spouse & Dependent premiums
  • Short- & Long-Term Disability
  • Group Term Life and AD&D
  • Voluntary Life and AD&D
  • 401(k) Plan
  • Equity Options
  • Flexible time off
  • MacBooks issued to all employees