DevOps/Site Reliability Engineer

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

DevOps/SRE engineer optimizing CI/CD, AWS infrastructure, and observability. Supporting technology tools for student athletes, coaches, and event operators.

Responsibilities:

  • Configure, manage, and improve Bitbucket pipelines for staging and production deployments
  • Improve CI pipeline speed, reliability, and security with the Cloud Security Engineer
  • Assist developers and QA teams with deployments
  • Work with Docker and AWS ECR for container builds and deployment workflows
  • Investigate system issues flagged by Sentry, New Relic, and CloudWatch
  • Monitor application performance, identify bottlenecks, and propose solutions
  • Respond to production and staging issues, including database latency, unresponsive resources, and failed jobs
  • Maintain and support non-production environments for developers and QA
  • Maintain and improve AWS infrastructure and Terraform resources
  • Update and upgrade AWS services to ensure reliability and scalability
  • Partner with engineers to design scalable, observable, and resilient systems
  • Ensure secure configurations in CI/CD, AWS, and containerized workloads
  • Contribute improvements to workflows, automation, and monitoring strategies
  • Leverage AI to automate monitoring and diagnosis

Requirements:

  • 3+ years of experience in DevOps, SRE, or related engineering roles
  • Strong experience configuring CI/CD pipelines, including Bitbucket Pipelines, GitHub Actions, or similar
  • Experience configuring, debugging, and deploying PHP applications
  • Hands-on experience with Docker and AWS ECR for container builds and deployments
  • Strong experience with AWS services such as EC2, RDS, ECS, and Lambda
  • Strong experience with Terraform for infrastructure as code
  • Familiarity with monitoring and observability tools such as New Relic, Sentry, and CloudWatch
  • Strong troubleshooting skills for performance issues in databases, applications, and distributed systems
  • Experience with modern software development workflows, including agile teams, code reviews, and branching strategies
  • Strong scripting and automation skills using Bash, Python, or similar
  • Excellent communication skills and a collaborative mindset
  • Interest in leveraging AI agents to automate monitoring and diagnosis workflows

Benefits:

  • Remote work arrangement
  • Opportunity to work with innovative tools involving mobile and web applications, computer vision, and LLMs
  • Collaboration with distributed teams across the United States