Site Reliability Engineer – Core Streaming

Posted 14hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer operating Yelp’s Kafka-based streaming infrastructure. Automating upgrades, scaling, incident recovery, and reliable data pipelines across Canada.

Responsibilities:

  • Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments
  • Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery
  • Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability
  • Troubleshoot complex issues affecting data flow, performance, or stability
  • Lead root cause analyses
  • Execute Kafka version upgrades and platform migrations with minimal disruption to critical services
  • Participate in on-call rotations using a geographically distributed follow-the-sun model
  • Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure

Requirements:

  • Solid SRE or infrastructure engineering foundation
  • Experience with infrastructure-as-code, especially Terraform
  • Experience with configuration management tools such as Puppet, Ansible, or equivalent
  • Experience with cloud platforms; AWS preferred
  • Linux operations experience
  • Production-level experience with Kafka or similar technologies at scale
  • Experience with cluster upgrades, migrations, and capacity planning
  • Programming proficiency in Python, Java, or similar
  • Strong debugging and systems-thinking skills across distributed systems
  • Experience with Apache Flink or other stream processing frameworks (nice to have)
  • Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams (nice to have)
  • Experience building internal self-service tooling or developer platforms (nice to have)
  • Experience with incident response and management (nice to have)

Benefits:

  • Fully remote work across Canada
  • Follow-the-sun on-call model; no one needs to be on-call 24 hours a day
  • Support from managers, mentors, and teams
  • Five star benefits (linked in the posting)
  • Reasonable accommodations for individuals with disabilities in the job application process