Technical Lead – GPU Infrastructure

Posted 5hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Technical Lead owning Tether Data’s GPU compute and managed inference platform architecture. Leading distributed engineers across Slurm, Kubernetes, bare-metal GPU infrastructure, and operations.

Responsibilities:

  • Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform
  • Lead architecture proposals, high-level and low-level designs, technical reviews, and maintain the architecture baseline
  • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation
  • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
  • Design, build, and operate a managed Slurm service for research users
  • Own Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation
  • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement
  • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
  • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers
  • Lead incident response, post-incident reviews, and design a sustainable on-call model
  • Serve as the primary technical interface to infrastructure partners and vendors
  • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
  • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
  • Complete the platform team and set the technical bar for new engineers

Requirements:

  • Eight or more years of hands-on engineering experience
  • At least three years leading teams that build and operate infrastructure platforms other teams depend on
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
  • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
  • Experience operating an HPC or GPU training cluster for a research population is ideally preferred
  • GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance
  • High-performance interconnect experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
  • Linux systems expertise including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads
  • Production Kubernetes operation, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design
  • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across many nodes
  • Observability and operations experience with Prometheus, Grafana and Loki or equivalents, SLOs, incident response, and post-incident review
  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services and make architecture decisions
  • Experience shipping a platform with real users, such as a multi-tenant IaaS or PaaS or research computing service
  • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives
  • Excellent written and spoken English
  • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers
  • Desirable experience with modern serving stacks such as vLLM, SGLang, or TensorRT-LLM
  • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships

Benefits:

  • Fully remote work arrangement
  • Occasional travel to partner sites and team events