Senior Solutions Architect – Large Scale Neural Networks Inference
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Solutions Architect optimizing large-scale neural network inference for NVIDIA’s EMEA AI customers. Architecting production pipelines and shaping the NVIDIA inference technology roadmap.
Responsibilities:
- Lead the inference strategy for a portfolio of EMEA AI Natives customers, guiding engagements from initial proof of concept to production-scale deployments
- Identify inference challenges across customer deployments, including latency, efficiency, cost per token, memory utilization, and low-latency networking
- Architect and optimize high-performance inference pipelines using NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other inference backends
- Improve GPU utilization and AI cluster efficiency
- Translate customer insights and deployment patterns into actionable product feedback
- Develop the roadmap for the NVIDIA stack, including Dynamo, TensorRT-LLM, and NIM
- Define the technical direction for AI inference across EMEA
- Align NVIDIA and customer organization stakeholders to influence strategic technology decisions for next-generation AI inference at scale
Requirements:
- MS or PhD in Computer Science, Engineering, or equivalent experience in the field
- 8+ years in AI/ML infrastructure, with deep expertise in LLM/VLM inference optimization and production deployment at scale
- Deep understanding of transformer inference acceleration: quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models
- Understanding of GPU memory hierarchies and low-latency networking and their influence on inference performance
- Proven track record to lead technical initiatives
- Excellent communication skills, effective with research scientists, infrastructure engineers, and executive team members
- Experience with NVIDIA's inference stack, including TensorRT-LLM, Triton Inference Server, NIM, and NVIDIA Dynamo
- Experience with GPU orchestration on Kubernetes
- Experience operating inference at scale inside a frontier AI lab or hyperscale's inference team
- Contributions to open-source inference projects such as vLLM, SGLang, KServe, or NVIDIA Dynamo
Benefits:
- Highly competitive salaries
- Comprehensive benefits package


















