Site Reliability Engineer – Core Streaming
Posted 14hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer operating Yelp’s Kafka-based streaming infrastructure. Automating upgrades, scaling, incident recovery, and reliable data pipelines across Canada.
Responsibilities:
- Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments
- Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery
- Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability
- Troubleshoot complex issues affecting data flow, performance, or stability
- Lead root cause analyses
- Execute Kafka version upgrades and platform migrations with minimal disruption to critical services
- Participate in on-call rotations using a geographically distributed follow-the-sun model
- Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure
Requirements:
- Solid SRE or infrastructure engineering foundation
- Experience with infrastructure-as-code, especially Terraform
- Experience with configuration management tools such as Puppet, Ansible, or equivalent
- Experience with cloud platforms; AWS preferred
- Linux operations experience
- Production-level experience with Kafka or similar technologies at scale
- Experience with cluster upgrades, migrations, and capacity planning
- Programming proficiency in Python, Java, or similar
- Strong debugging and systems-thinking skills across distributed systems
- Experience with Apache Flink or other stream processing frameworks (nice to have)
- Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams (nice to have)
- Experience building internal self-service tooling or developer platforms (nice to have)
- Experience with incident response and management (nice to have)
Benefits:
- Fully remote work across Canada
- Follow-the-sun on-call model; no one needs to be on-call 24 hours a day
- Support from managers, mentors, and teams
- Five star benefits (linked in the posting)
- Reasonable accommodations for individuals with disabilities in the job application process











