Senior Site Reliability Engineer

Posted 23hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior Site Reliability Engineer maintaining Akamai's Compute services and infrastructure. Improving reliability, automation, observability, and incident response for the distributed cloud and edge platform.

Responsibilities:

  • Provide support and mentorship for other SRE engineers
  • Define requirements throughout the product lifecycle and influence new designs and standards
  • Deploy and maintain internal platforms and tools
  • Partner with multiple teams to ensure product and service availability, reliability, scalability, and usability
  • Improve the Compute Cloud Interface platform for faster error detection and remediation, performance, and reliability
  • Develop and improve automation to support daily activities and reduce toil
  • Participate in on-call rotations and guide restoration and repair of service-impacting issues
  • Troubleshoot and resolve customer escalations and incidents with internal teams

Requirements:

  • 5 years of relevant experience
  • Bachelor's degree in Computer Science or its equivalent
  • Experience automating with Python and/or Golang and bash scripting
  • Understanding of systems reliability, observability, monitoring, and adherence to SLOs
  • Experience with SaltStack, Terraform, Ansible, and Jenkins CI/CD
  • Hands-on mastery of Linux administration
  • Experience with container-based platforms such as Docker
  • Experience with Prometheus, Grafana, Loki, nginx, Envoy, HAProxy, and Redis

Benefits:

  • Benefits supporting health, well-being, finances, and life beyond work
  • FlexBase flexible workplace options: work at home, in an office, or a combination of both