Senior Software Engineer – AI Inference & Runtime Platform
Posted 4hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Software Engineer building AI inference infrastructure and secure agent sandboxes at AZX. Operating Kubernetes, GPU serving, microVM isolation, and stateful data platforms.
Responsibilities:
- Own the inference control plane for open-weight models and custom task-model zoos across managed GPU clouds and customer-managed Kubernetes clusters
- Manage serving-tier engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline
- Administer the Kubernetes layer for inference and sandbox workloads, including operators, CRDs, autoscaling, GPU scheduling and sharing, and node lifecycle
- Own the stateful data plane for hosted vector stores and graph stores, including deployment, backups, scaling, recovery, and tested restores
- Oversee the sandbox runtime and host-side control plane, including lifecycle, exec, snapshot/fork, teardown, metering, and isolation-boundary threat modeling
- Direct FastAPI control-plane services, Terraform/OpenTofu, Bicep, and dashboards
- Run the platform layer supporting backend services and debug into backend services as needed
- Manage the open-source posture through reviewed PRs, documentation, and reproducible builds
- Build the fork engine, guest agent, and multi-substrate model lifecycle using Rust, Kubernetes controllers, and FastAPI endpoints
- Account for per-token costs accurately, including client mid-stream disconnections
- Operate hardware-isolated microVMs for secure, compliant execution of untrusted agent-generated code
Requirements:
- 5+ years of shipping production systems in a systems language
- Rust is the house language; deep Go, C/C++, or Zig with genuine appetite for Rust counts
- Familiarity with async runtimes, memory-safety discipline, and debugging at the syscall boundary
- Experience operating Kubernetes workloads, including controllers or operators, scheduling, autoscaling, and node lifecycle
- Ability to threat-model isolation boundaries, including namespaces, cgroups, seccomp, and hypervisors
- Experience applying security best practices for agentic execution: least privilege, no credentials in the sandbox, audit trails, and human approval on write actions
- Hands-on experience deploying or operating open-weight LLM serving infrastructure such as vLLM/SGLang or similar
- Performance discipline in distributed systems
- Practical depth in Rust (tokio), Python/FastAPI, Kubernetes operators, KEDA, Karpenter, GPU device plugins/DRA
- Familiarity with Firecracker, Kata, gVisor, or comparable isolation technology
- Familiarity with secrets management and egress control, such as Vault/KMS-class tools
- Experience hosting stateful systems including vector stores and graph stores, with backup and failover discipline
- Comfort operating across AWS/Azure/GCP and managed GPU clouds
- Experience with Terraform/OpenTofu, Bicep, and OpenTelemetry
- Bachelor's Degree; Master's is a plus
- Must be currently authorized to work in the United States on a full-time basis
- Must be able to travel 2x/year for company summits
Benefits:
- Competitive early-stage startup compensation (based on capabilities, experience, and location)
- Bonus eligibility
- Health insurance with meaningful coverage for dependents
- Flexible paid time off
- Equity
- Fully remote culture with a cluster of teammates in Seattle
- Company summits 2x/year


















