Site Reliability Engineer – Core Streaming
Posted 22hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer operating Yelp’s Kafka and Flink streaming platform across Canada. Automating cluster management, scaling, upgrades, migrations, and incident recovery for real-time data systems.
Responsibilities:
- Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments
- Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery
- Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability
- Troubleshoot complex issues affecting data flow, performance, or stability, and lead root cause analyses
- Execute Kafka version upgrades and platform migrations with minimal disruption to critical services
- Participate in on-call rotations
- Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure
Requirements:
- Solid SRE or infrastructure engineering foundation
- Experience with infrastructure-as-code, especially Terraform
- Experience with configuration management tools such as Puppet, Ansible, or equivalent
- Experience with cloud platforms; AWS preferred
- Linux operations experience
- Production-level experience with Kafka or similar technologies at scale
- Experience with cluster upgrades, migrations, and capacity planning
- Programming proficiency in Python, Java, or similar
- Strong debugging and systems-thinking skills
- Comfortable tracing data-flow issues end-to-end across distributed systems
- Experience with Apache Flink or other stream-processing frameworks
- Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams
- Experience building internal self-service tooling or developer platforms
- Experience with incident response and management
Benefits:
- Fully remote work across Canada
- Yelp's five star benefits
- Support from managers, mentors, and teams
- Follow-the-sun on-call model, so no one needs to be on-call 24 hours a day
- Reasonable accommodations for individuals with disabilities
- Equal opportunity employment
















