Staff Platform Engineer
Posted 3hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff Platform Engineer defining cloud and DevSecOps strategy for Robots & Pencils’ enterprise AI systems. Leading Kubernetes, AI/ML infrastructure, migrations, reliability, security, and platform standards remotely in Canada.
Responsibilities:
- Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems
- Architect and own scalable Kubernetes platforms and containerized infrastructure at scale
- Own infrastructure as code strategy and standards across environments
- Lead DevSecOps implementation, including secrets management, compliance, auditing, IAM, and zero-trust networking
- Drive platform reliability, performance SLAs, and cost optimization across production systems
- Lead complex cloud migrations and platform modernization initiatives
- Own observability strategy and production reliability practices
- Lead AI/ML platform infrastructure design and operation, including model serving and deployment, GPU workload orchestration, LLM gateway and observability, vector store infrastructure, and CI/CD for AI/ML systems
- Use modern AI assistants such as Claude and Cursor to improve delivery quality and pace
- Partner with engineering, product, and leadership to align platform strategy with business and delivery goals
- Communicate infrastructure decisions and tradeoffs to technical and non-technical stakeholders
- Lead design reviews, architecture discussions, and release readiness assessments
- Establish platform engineering standards and best practices
- Mentor junior and mid-level engineers
- Act as a technical escalation point for complex infrastructure and platform challenges
- Evaluate emerging tools and technologies to improve platform reliability and developer experience
Requirements:
- 7+ years of professional DevOps or platform engineering experience, including leading complex platform initiatives
- Expert scripting and programming skills, such as Python, Go, Java, or Bash
- Deep expertise in at least one major cloud platform
- Expert Kubernetes and container orchestration skills
- Expert infrastructure-as-code skills across multiple tools
- Strong CI/CD architecture experience at scale
- Strong DevSecOps experience, including secrets management, compliance, and auditing
- Experience with networking, IAM, cloud security architecture, and zero-trust principles
- Experience with service mesh, distributed systems, and microservices architecture
- Strong AI/ML platform infrastructure experience, including model serving and deployment, GPU workload orchestration, LLM gateways and observability, vector store infrastructure, and CI/CD for AI/ML systems
- Demonstrated leadership and technical mentoring experience
- Strong stakeholder communication skills
- Demonstrable day-to-day usage and expert knowledge of AI-forward tools such as Claude and Cursor
- Excellent problem-solving skills and sound judgment in ambiguous technical and business challenges
- Cloud certifications or FinOps experience is a plus
- Helpful: HPC cluster infrastructure across AWS, CoreWeave, GCP, and OCI
- Helpful: HPC job schedulers and workload managers such as Slurm or equivalent
















