Senior Solutions Architect – Large Scale AI Inference

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior Solutions Architect guiding EMEA customers deploying NVIDIA’s large-scale AI inference on GPU clusters. Optimizing MoE serving, interconnect-aware scheduling, and next-generation inference architectures.

Responsibilities:

  • Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters
  • Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs
  • Improve inference efficiency across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments
  • Collaborate with NVIDIA product teams, including Dynamo, TensorRT-LLM, and NIXL, to accelerate customer success
  • Animate the AI inference developer community across EMEA through technical workshops, hackathons, and reference architectures
  • Establish technical direction for scalable, high-performance AI inference across demanding production environments

Requirements:

  • MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience
  • 5+ years of experience in Neural Networks inference optimization
  • Solid understanding of transformers inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache optimization
  • Practical experience in MoE inference at scale, including expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale
  • Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level
  • Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling
  • Understanding of GPU memory hierarchies and high-speed interconnects, including NVLink, InfiniBand, RDMA, and UCX
  • Contributions to advanced AI labs or large-scale AI infrastructure providers performing inference on thousands of GPUs
  • Published work or benchmarks in large-scale AI inference

Benefits:

  • Highly competitive salaries
  • Comprehensive benefits package