Site Reliability Engineer

Posted 2ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer operating scalable cloud infrastructure for a client's Cloud Operations team. Automating deployments, monitoring, incident response, and Kubernetes migration for high-concurrency production systems.

Responsibilities:

  • Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents
  • Use on-call shifts to prevent incidents from recurring
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes
  • Configure monitoring and alerting to detect symptoms rather than outages
  • Document every action so findings become repeatable actions and automation
  • Improve the deployment process
  • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
  • Debug production issues across services and stack levels
  • Plan infrastructure growth
  • Code infrastructure automation with Ansible and Terraform
  • Improve Prometheus monitoring or build new metrics
  • Help release managers deploy and fix new application software versions
  • Plan and execute migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
  • Develop relationships with product groups and define their SRE KPIs

Requirements:

  • Think cloud-first regardless of public cloud provider
  • Think security-first
  • Understand systems, including edge cases, failure modes, behaviors, and specific implementations
  • Know Linux and Windows
  • Know configuration-management systems such as Ansible or Puppet
  • Strong programming skills in Python, Java, Golang, or Node.js
  • Collaborate and communicate asynchronously
  • Document work thoroughly
  • Have a go-for-it attitude and fix broken systems
  • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies