Technical Support Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Technical Support Engineer troubleshooting Linux, networking, storage, and GPU issues for Hyperbolic Labs' AI cloud. Owning customer tickets, SLAs, runbooks, and critical on-call escalations.

Responsibilities:

  • Own customer support tickets end to end, from first response through resolution, including after escalation
  • Manage the SLA clock, including first response, severity classification, and response commitments
  • Reproduce customer problems and gather relevant logs and configuration
  • Determine whether faults originate with Hyperbolic or an infrastructure provider
  • Troubleshoot customer environment access and configuration, including SSH keys, NFS mounts and storage, quotas, security groups, containers, drivers, billing, and accounts
  • Execute documented runbooks and author new runbooks for novel solutions
  • Keep customer-facing documentation and the internal knowledge base current
  • Collaborate with other engineers while retaining ownership of escalated issues
  • Participate in an on-call rotation for critical issues
  • Support customers using Hyperbolic Labs' Open-Access AI Cloud, GPU marketplace, and AI inference service

Requirements:

  • Very strong Linux experience and daily work in the CLI
  • Experience owning tickets against a response SLA in cloud, hosting, or infrastructure support
  • Solid networking and storage fundamentals: SSH, NFS and mounts, DNS, firewalls and security groups
  • Working familiarity with GPU workloads: nvidia-smi, drivers, CUDA, containers
  • Clear and fast written communication under time pressure
  • Good judgment about the limits of your own knowledge, with a bias toward escalating early with a complete picture
  • Comfortable working across time zones and with an on-call rotation for critical issues
  • Experience with ticketing and on-call tooling such as Zendesk, Linear, PagerDuty, or similar
  • Scripting in bash or Python to automate repeat work
  • Exposure to Slurm, Kubernetes, or Docker in a multi-tenant environment
  • Background in GPU cloud, HPC, or a hardware-adjacent support organization