Staff Production Engineer
Posted 11ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff Production Engineer improving reliability, performance, and resilience across Canva’s large-scale design platform. Embedding with product and infrastructure teams to solve distributed-systems risks.
Responsibilities:
- Own a long-term engagement area within one of Canva's highest-risk technical domains
- Write production software to improve reliability, efficiency, and resilience
- Instrument, refactor, and rebuild software components that cause problems at scale
- Work embedded alongside product and infrastructure teams
- Pair with, mentor, and learn from fellow production engineers
- Reduce incidents, accelerate recovery, lower incident severity, and improve latency
- Contribute shared platform capabilities when recurring patterns emerge
- Build system-level leverage by improving how engineers develop on top of critical systems
- Develop trusted relationships with embedded teams and guide them toward shipping faster with more confidence and less toil
Requirements:
- Experience owning reliability work within large-scale distributed systems
- Experience working as an engineer embedded in or closely partnering with a product or feature team
- Production-scale experience with Java, Go, Rust, C++, or a comparable systems language
- Practical experience with sharding, replication, failure modes, and consistency tradeoffs
- Ability to debug large, unfamiliar codebases
- Proven influence without authority through technical expertise and trust
- Networking depth
- Linux internals and kernel-level understanding
- Knowledge of distributed systems patterns including consistent hashing, leader election, consensus, backpressure, and circuit breakers
- Experience with observability tooling, tracing, dashboards, alerting, and SLOs
- Kubernetes and production-scale container orchestration experience
- Performance profiling of JVM applications or systems-level processes
- Meaningful AWS cloud infrastructure experience
- Serious production-environment on-call and incident-response experience
- Enterprise SaaS background (nice to have)
- JVM internals experience, including GC tuning and thread profiling (nice to have)
- Multi-region or sharding experience (nice to have)
- eBPF or kernel instrumentation experience (nice to have)
Benefits:
- Equity packages
- Inclusive parental leave policy that supports all parents & carers
- Annual Vibe & Thrive allowance supporting wellbeing, social connection, office setup & more
- Flexible leave options
- Flexible remote work arrangements







