DevOps/Site Reliability Engineer
Posted 4hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
DevOps/SRE engineer optimizing CI/CD, AWS infrastructure, and observability. Supporting technology tools for student athletes, coaches, and event operators.
Responsibilities:
- Configure, manage, and improve Bitbucket pipelines for staging and production deployments
- Improve CI pipeline speed, reliability, and security with the Cloud Security Engineer
- Assist developers and QA teams with deployments
- Work with Docker and AWS ECR for container builds and deployment workflows
- Investigate system issues flagged by Sentry, New Relic, and CloudWatch
- Monitor application performance, identify bottlenecks, and propose solutions
- Respond to production and staging issues, including database latency, unresponsive resources, and failed jobs
- Maintain and support non-production environments for developers and QA
- Maintain and improve AWS infrastructure and Terraform resources
- Update and upgrade AWS services to ensure reliability and scalability
- Partner with engineers to design scalable, observable, and resilient systems
- Ensure secure configurations in CI/CD, AWS, and containerized workloads
- Contribute improvements to workflows, automation, and monitoring strategies
- Leverage AI to automate monitoring and diagnosis
Requirements:
- 3+ years of experience in DevOps, SRE, or related engineering roles
- Strong experience configuring CI/CD pipelines, including Bitbucket Pipelines, GitHub Actions, or similar
- Experience configuring, debugging, and deploying PHP applications
- Hands-on experience with Docker and AWS ECR for container builds and deployments
- Strong experience with AWS services such as EC2, RDS, ECS, and Lambda
- Strong experience with Terraform for infrastructure as code
- Familiarity with monitoring and observability tools such as New Relic, Sentry, and CloudWatch
- Strong troubleshooting skills for performance issues in databases, applications, and distributed systems
- Experience with modern software development workflows, including agile teams, code reviews, and branching strategies
- Strong scripting and automation skills using Bash, Python, or similar
- Excellent communication skills and a collaborative mindset
- Interest in leveraging AI agents to automate monitoring and diagnosis workflows
Benefits:
- Remote work arrangement
- Opportunity to work with innovative tools involving mobile and web applications, computer vision, and LLMs
- Collaboration with distributed teams across the United States















