Staff Site Reliability Operations Engineer

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.

Responsibilities:

  • Lead global platform reliability and drive the next-generation observability strategy on Google Cloud Platform
  • Architect, optimize, and troubleshoot networking infrastructure across OSI Layers 1–7
  • Design, scale, and optimize the Grafana Labs observability stack, including Grafana, Mimir, Loki, Tempo, and Beyla
  • Deploy machine-learning models and automated anomaly detection to reduce alert fatigue and predict bottlenecks
  • Architect, scale, secure, and manage production Google Kubernetes Engine clusters
  • Tune and maintain high-throughput Apache Kafka clusters for low-latency event delivery and high availability
  • Ensure performance, scalability, and disaster-recovery readiness across PostgreSQL, AlloyDB, and BigQuery
  • Automate incident triage, root-cause analysis, and remediation through Grafana workflows and AIOps insights
  • Champion the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards
  • Coach engineers on advanced debugging, distributed-systems thinking, and intelligent operations

Requirements:

  • Proven track record of high autonomy and successful delivery in a 100% remote engineering environment
  • 8+ years of experience in SRE, Production Engineering, or Distributed Systems infrastructure roles
  • Deep technical knowledge across OSI Layers 1–7
  • Physical/fiber infrastructure awareness, switching, BGP, and OSPF
  • TCP congestion control, UDP, and QUIC
  • Session management, TLS termination, DNS architecture, HTTP/3, and gRPC
  • Expert-level mastery of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows
  • Experience managing high-throughput Apache Kafka pipelines and large-scale PostgreSQL, AlloyDB, and BigQuery environments
  • Hands-on experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
  • Experience applying AI/ML to time-series anomaly detection, log clustering, and correlation
  • Advanced production-scale expertise with HashiCorp Terraform for multi-region GCP architectures
  • High proficiency in Go and Python
  • Exceptional written and verbal communication skills for asynchronous alignment
  • Deep knowledge of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization
  • Understanding of Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump

Benefits:

  • Eligibility for a bonus as part of the total compensation package
  • Benefits package available; details provided via the employer's benefits information