Site Reliability Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer improving CloudBlue’s multi-tenant SaaS reliability, scalability, and observability. Operating Kubernetes platforms, leading incident response, and strengthening cloud infrastructure for service providers worldwide.

Responsibilities:

  • Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services
  • Influence system architecture with a focus on reliability, scalability, and operability
  • Reduce operational toil through automation and process improvement
  • Design and operate the observability stack across metrics, logs, and traces
  • Develop alerting strategies and dashboards for platform and business health
  • Design and maintain high-availability architectures, redundancy, failover, and disaster recovery strategies
  • Conduct capacity planning, load testing, and performance optimization
  • Lead production incident coordination, communication, and service restoration
  • Own blameless postmortems and drive improvements to reduce incidents, MTTR, and customer impact
  • Improve reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing
  • Partner with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability
  • Maintain runbooks and operational documentation
  • Promote SRE best practices across engineering teams
  • Support other tasks or projects as assigned

Requirements:

  • 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer
  • Strong ownership of production systems
  • Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
  • Hands-on experience with Datadog, Grafana, and Elasticsearch/Kibana
  • Solid understanding of Linux, networking, and distributed systems fundamentals
  • Experience with Docker and Kubernetes
  • Strong scripting and automation skills using Python and/or Bash
  • Experience participating in on-call rotations and production incident response
  • Strong written and spoken English
  • Cloud experience, preferably with Azure; AWS and/or GCP experience valued
  • Experience with hybrid or on-premises integrations beneficial
  • Experience defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing are advantageous or considered assets

Benefits:

  • A competitive salary that values you and your unique skill sets
  • Career advancement & professional development opportunities
  • Flexible work arrangements to support work/life balance
  • 24/7 award-winning customer support
  • Diversity and inclusion
  • Accommodation may be provided in all parts of the hiring process