Site Reliability Engineer
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer improving CloudBlue’s multi-tenant SaaS reliability, scalability, and observability. Operating Kubernetes platforms, leading incident response, and strengthening cloud infrastructure for service providers worldwide.
Responsibilities:
- Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services
- Influence system architecture with a focus on reliability, scalability, and operability
- Reduce operational toil through automation and process improvement
- Design and operate the observability stack across metrics, logs, and traces
- Develop alerting strategies and dashboards for platform and business health
- Design and maintain high-availability architectures, redundancy, failover, and disaster recovery strategies
- Conduct capacity planning, load testing, and performance optimization
- Lead production incident coordination, communication, and service restoration
- Own blameless postmortems and drive improvements to reduce incidents, MTTR, and customer impact
- Improve reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing
- Partner with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability
- Maintain runbooks and operational documentation
- Promote SRE best practices across engineering teams
- Support other tasks or projects as assigned
Requirements:
- 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer
- Strong ownership of production systems
- Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
- Hands-on experience with Datadog, Grafana, and Elasticsearch/Kibana
- Solid understanding of Linux, networking, and distributed systems fundamentals
- Experience with Docker and Kubernetes
- Strong scripting and automation skills using Python and/or Bash
- Experience participating in on-call rotations and production incident response
- Strong written and spoken English
- Cloud experience, preferably with Azure; AWS and/or GCP experience valued
- Experience with hybrid or on-premises integrations beneficial
- Experience defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing are advantageous or considered assets
Benefits:
- A competitive salary that values you and your unique skill sets
- Career advancement & professional development opportunities
- Flexible work arrangements to support work/life balance
- 24/7 award-winning customer support
- Diversity and inclusion
- Accommodation may be provided in all parts of the hiring process











