Senior Site Reliability Engineer
Posted 23hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Site Reliability Engineer maintaining Akamai's Compute services and infrastructure. Improving reliability, automation, observability, and incident response for the distributed cloud and edge platform.
Responsibilities:
- Provide support and mentorship for other SRE engineers
- Define requirements throughout the product lifecycle and influence new designs and standards
- Deploy and maintain internal platforms and tools
- Partner with multiple teams to ensure product and service availability, reliability, scalability, and usability
- Improve the Compute Cloud Interface platform for faster error detection and remediation, performance, and reliability
- Develop and improve automation to support daily activities and reduce toil
- Participate in on-call rotations and guide restoration and repair of service-impacting issues
- Troubleshoot and resolve customer escalations and incidents with internal teams
Requirements:
- 5 years of relevant experience
- Bachelor's degree in Computer Science or its equivalent
- Experience automating with Python and/or Golang and bash scripting
- Understanding of systems reliability, observability, monitoring, and adherence to SLOs
- Experience with SaltStack, Terraform, Ansible, and Jenkins CI/CD
- Hands-on mastery of Linux administration
- Experience with container-based platforms such as Docker
- Experience with Prometheus, Grafana, Loki, nginx, Envoy, HAProxy, and Redis
Benefits:
- Benefits supporting health, well-being, finances, and life beyond work
- FlexBase flexible workplace options: work at home, in an office, or a combination of both












