Site Reliability Engineer
Posted 2ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer operating scalable cloud infrastructure for a client's Cloud Operations team. Automating deployments, monitoring, incident response, and Kubernetes migration for high-concurrency production systems.
Responsibilities:
- Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents
- Use on-call shifts to prevent incidents from recurring
- Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes
- Configure monitoring and alerting to detect symptoms rather than outages
- Document every action so findings become repeatable actions and automation
- Improve the deployment process
- Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
- Debug production issues across services and stack levels
- Plan infrastructure growth
- Code infrastructure automation with Ansible and Terraform
- Improve Prometheus monitoring or build new metrics
- Help release managers deploy and fix new application software versions
- Plan and execute migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
- Develop relationships with product groups and define their SRE KPIs
Requirements:
- Think cloud-first regardless of public cloud provider
- Think security-first
- Understand systems, including edge cases, failure modes, behaviors, and specific implementations
- Know Linux and Windows
- Know configuration-management systems such as Ansible or Puppet
- Strong programming skills in Python, Java, Golang, or Node.js
- Collaborate and communicate asynchronously
- Document work thoroughly
- Have a go-for-it attitude and fix broken systems
- Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies


















