Site Reliability Engineer
Posted 4hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer keeping user-facing services and production systems reliable for a client's cloud operations team. Automating infrastructure, improving monitoring, and scaling Kubernetes-based deployments.
Responsibilities:
- Participate in an on-call rotation responding to production availability incidents and support service engineers with customer incidents
- Use on-call shifts to prevent incidents from recurring
- Operate infrastructure with Ansible, Puppet, Terraform, and Kubernetes
- Build monitoring and alerting focused on symptoms rather than outages
- Document actions and convert findings into repeatable procedures and automation
- Improve the deployment process
- Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
- Debug production issues across services and stack levels
- Plan infrastructure growth
- Code infrastructure automation with Ansible and Terraform
- Improve Prometheus monitoring and build new metrics
- Help release managers deploy and fix new application versions
- Plan and execute migration from AWS virtual machines to cloud-native, container-based Kubernetes deployments on EKS
- Develop relationships with product groups and define SRE KPIs
Requirements:
- Think cloud-first across public cloud environments
- Think security-first
- Understand systems, edge cases, failure modes, behaviors, and specific implementations
- Familiarity with Linux and Windows
- Knowledge of configuration-management systems such as Ansible or Puppet
- Strong programming skills in Python, Java, Golang, or Node.js
- Ability to collaborate and communicate asynchronously and document work
- Go-for-it attitude and willingness to fix broken systems
- Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies
















