Site Reliability Engineer

Posted 2ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer operating scalable cloud infrastructure and production services for a client. Automating deployments, monitoring, incident response, and AWS-to-Kubernetes migrations.

Responsibilities:

  • Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents
  • Use on-call shifts to prevent incidents from recurring
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes
  • Configure monitoring and alerting to report symptoms rather than outages
  • Document every action so findings become repeatable actions and automation
  • Improve the deployment process
  • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
  • Debug production issues across services and stack levels
  • Plan infrastructure growth
  • Code infrastructure automation with Ansible and Terraform
  • Improve Prometheus monitoring or build new metrics
  • Help release managers deploy and fix new application software versions
  • Plan and execute migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
  • Develop relationships with product groups and define SRE KPIs

Requirements:

  • Think cloud-first regardless of public-cloud provider
  • Think security-first
  • Understand systems, including edge cases, failure modes, behaviours, and specific implementations
  • Knowledge of Linux and Windows
  • Knowledge of configuration-management systems such as Ansible or Puppet
  • Strong programming skills in Python, Java, Golang, or Node.js
  • Ability to collaborate and communicate asynchronously and document work thoroughly
  • Go-for-it attitude and willingness to fix broken systems
  • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies