Senior Solutions Architect – Large Scale AI Inference
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Solutions Architect guiding EMEA customers deploying NVIDIA’s large-scale AI inference on GPU clusters. Optimizing MoE serving, interconnect-aware scheduling, and next-generation inference architectures.
Responsibilities:
- Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters
- Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs
- Improve inference efficiency across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments
- Collaborate with NVIDIA product teams, including Dynamo, TensorRT-LLM, and NIXL, to accelerate customer success
- Animate the AI inference developer community across EMEA through technical workshops, hackathons, and reference architectures
- Establish technical direction for scalable, high-performance AI inference across demanding production environments
Requirements:
- MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience
- 5+ years of experience in Neural Networks inference optimization
- Solid understanding of transformers inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache optimization
- Practical experience in MoE inference at scale, including expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale
- Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level
- Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling
- Understanding of GPU memory hierarchies and high-speed interconnects, including NVLink, InfiniBand, RDMA, and UCX
- Contributions to advanced AI labs or large-scale AI infrastructure providers performing inference on thousands of GPUs
- Published work or benchmarks in large-scale AI inference
Benefits:
- Highly competitive salaries
- Comprehensive benefits package


















