Senior Site Reliability Engineer
Posted 2ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Site Reliability Engineer managing GCP infrastructure for an AI/ML geospatial intelligence company. Improving observability, incident response, reliability metrics, and cloud cost governance.
Responsibilities:
- Design and evolve cloud infrastructure on GCP at scale
- Build internal tooling and automation that promote team autonomy and self-service
- Advance the observability platform using metrics, logging, and tracing to reduce MTTR
- Build visibility into infrastructure costs and drive governance and optimization initiatives
- Champion reliability best practices including SLOs, SLIs, error budgets, and DORA metrics
- Lead incident management, facilitate post-incident reviews, and participate in an on-call rotation
- Work cross-functionally to optimize cloud usage and cost
Requirements:
- 3+ years of Site Reliability Engineering or production SRE experience
- Proficiency with Google Cloud Platform (GCP), including cost optimization and governance
- Hands-on experience with Kubernetes for cluster and workload management
- Infrastructure as Code experience using Terraform or Deployment Manager
- Scripting and automation skills in Python, Bash, or Go
- Strong observability stack experience with Prometheus, Grafana, OpenTelemetry, logging, and distributed tracing
- Experience with incident management, post-incident reviews, and on-call rotation
- Ability to define and implement SLOs, SLIs, and error budgets
- Familiarity with DORA metrics is nice to have
- Background in AI/ML or geospatial technology environments is nice to have
- Must have work authorization in the applicable country
- Visa sponsorship is not available
Benefits:
- Fully remote work
- Compensation commensurate with experience and location
- Visa sponsorship is not available

















