Senior Site Reliability Engineer

Posted 9hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior SRE auditing and scaling infrastructure for Texas Sports Academy, an AI-first K-12 school. Improving AWS reliability, observability, deployments, and incident response.

Responsibilities:

  • Audit infrastructure, deployment pipelines, monitoring, alerting, and incident response end to end
  • Identify brittleness, scaling risks, and unnecessary costs
  • Write clear audit reports explaining findings, impact, and fixes
  • Implement code, configuration, monitoring, alerting, and deployment improvements
  • Mature on-call, incident, SLO, and postmortem practices
  • Advise on infrastructure architecture decisions for scale
  • Partner directly with engineers through pairing, reviews, and clean handoffs

Requirements:

  • Eight or more years of hands-on site reliability, infrastructure, or production engineering work
  • Deep AWS cloud infrastructure expertise, including networking, IAM, VPCs, and failure modes
  • Experience building monitoring, alerting, and observability from the ground up with Datadog, Grafana, Prometheus, or comparable tools
  • Experience building and maintaining CI/CD and deployment pipelines
  • Comfortable writing production code and configuration and taking responsibility for it
  • Daily use of AI in workflows and comfort using AI coding tools
  • Strong written communication skills for producing actionable audit reports
  • Fully remote, open globally
  • Reliable internet and a quiet workspace
  • On-call and incident response leadership experience (bonus)
  • Cloud cost optimization experience (bonus)
  • Security-adjacent work experience (bonus)

Benefits:

  • Part-time consulting contract
  • Ongoing duration
  • Possibility of moving into a full-time role based on fit and results
  • Flexible hours based on scope
  • Fully remote, anywhere in the world
  • Opportunity to make decisions that keep systems running as the company scales
  • Fast deployment of changes
  • Equal opportunity employer