DevOps / Site Reliability Engineer

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

DevOps/SRE improving AWS infrastructure, CI/CD, and application reliability. Supporting technology tools for student athletes, coaches, and event operators.

Responsibilities:

  • Configure, manage, and improve Bitbucket pipelines for deploying applications to staging and production
  • Improve CI pipeline speed, reliability, and security in collaboration with the Cloud Security Engineer
  • Assist developers and QA teams with deployments
  • Work with Docker and AWS ECR for container builds and deployment workflows
  • Review and investigate system issues flagged by Sentry, New Relic, and CloudWatch
  • Monitor application performance, identify bottlenecks, and propose solutions
  • Respond to production and staging issues, including database latency, unresponsive resources, or failed jobs
  • Maintain and support non-production environments used by developers and QA
  • Maintain and improve AWS infrastructure and Terraform resources
  • Perform updates and upgrades to AWS services to ensure reliability and scalability
  • Partner with engineers to design scalable, observable, and resilient systems
  • Work with the cloud security engineer to ensure secure configurations in CI/CD, AWS, and containerized workloads
  • Contribute improvements to workflows, automation, and monitoring strategies
  • Leverage AI to automate monitoring and diagnosis

Requirements:

  • 3+ years of experience in DevOps, SRE, or related engineering roles
  • Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar)
  • Experience configuring, debugging and deploying PHP applications
  • Hands-on experience with Docker and AWS ECR for container builds and deployments
  • Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code
  • Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar
  • Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems
  • Experience with modern software development workflows (agile teams, code reviews, branching strategies)
  • Strong scripting and automation skills (Bash, Python, or similar)
  • Excellent communication skills and a collaborative mindset
  • Interest in leveraging AI agents to automate monitoring and diagnosis workflows