DevOps / Site Reliability Engineer
Posted 4hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
DevOps/SRE improving AWS infrastructure, CI/CD, and application reliability. Supporting technology tools for student athletes, coaches, and event operators.
Responsibilities:
- Configure, manage, and improve Bitbucket pipelines for deploying applications to staging and production
- Improve CI pipeline speed, reliability, and security in collaboration with the Cloud Security Engineer
- Assist developers and QA teams with deployments
- Work with Docker and AWS ECR for container builds and deployment workflows
- Review and investigate system issues flagged by Sentry, New Relic, and CloudWatch
- Monitor application performance, identify bottlenecks, and propose solutions
- Respond to production and staging issues, including database latency, unresponsive resources, or failed jobs
- Maintain and support non-production environments used by developers and QA
- Maintain and improve AWS infrastructure and Terraform resources
- Perform updates and upgrades to AWS services to ensure reliability and scalability
- Partner with engineers to design scalable, observable, and resilient systems
- Work with the cloud security engineer to ensure secure configurations in CI/CD, AWS, and containerized workloads
- Contribute improvements to workflows, automation, and monitoring strategies
- Leverage AI to automate monitoring and diagnosis
Requirements:
- 3+ years of experience in DevOps, SRE, or related engineering roles
- Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar)
- Experience configuring, debugging and deploying PHP applications
- Hands-on experience with Docker and AWS ECR for container builds and deployments
- Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code
- Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar
- Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems
- Experience with modern software development workflows (agile teams, code reviews, branching strategies)
- Strong scripting and automation skills (Bash, Python, or similar)
- Excellent communication skills and a collaborative mindset
- Interest in leveraging AI agents to automate monitoring and diagnosis workflows















