Solutions Architect, Networking

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Solutions Architect designing and deploying NVIDIA GPU networking infrastructure for Canadian cloud partners. Optimizing large-scale AI training and inference systems across customer environments.

Responsibilities:

  • Become the trusted technical advisor for NVIDIA Cloud Partners in Canada to bring NVIDIA Data Center GPU and networking platforms to market at scale
  • Collaborate directly with customers to build, deploy, and optimize large-scale AI training and inference infrastructure using NVIDIA technology
  • Analyze deployment and performance data, identifying product health trends, system bottlenecks, and operational risks
  • Solve technical problems involving GPUs, networking, drivers, containers, firmware, and distributed system interactions
  • Deliver executive-level communication on status, risks, progress, and required decisions
  • Collaborate with internal engineering, product, and business teams on performance analysis and modeling of large GPU clusters
  • Travel to customer sites and industry events occasionally, up to 20%

Requirements:

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering, or related fields, or equivalent experience
  • 5+ years of Solution Architecture or similar Sales Engineering, Systems Engineering, Cloud Engineering, or Solution Engineering experience
  • Understanding of high-performance networking technologies, including RDMA, congestion control, and high-bandwidth interconnects
  • Hands-on experience bringing up and validating large-scale NVIDIA GPU platforms, including multi-GPU and multi-node architectures
  • Familiarity with NVIDIA system software stacks, including CUDA, NCCL, NVSwitch/NVLink, driver behavior, and performance tuning
  • Ability to identify performance bottlenecks at the cluster, node, accelerator, network, or application layer
  • Strong Linux fundamentals across drivers, kernel subsystems, cgroups, containers, and node-level performance analysis
  • Excellent presentation, communication, and collaboration skills
  • Prior experience deploying or optimizing deep learning training and inference at scale in production environments is advantageous
  • Familiarity with NVIDIA hardware and systems technology such as GPUs, networking, storage, NCCL, DCGM, UFM, Mission Control, and Base Command Manager is advantageous
  • Demonstrated leadership resolving multi-team infrastructure challenges is advantageous
  • Record of taking GPU or infrastructure products from pilot to high-volume deployment is advantageous
  • Ability to travel to customer sites and industry events occasionally, up to 20%

Benefits:

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity
  • Benefits for employees and families
  • Remote work arrangement
  • Occasional travel to customer sites and industry events