Senior Solutions Architect – Large Scale AI Inference
This role involves guiding EMEA-based AI companies and enterprises in deploying and optimizing large-scale AI inference workloads on multi-node GPU clusters. The Senior Solutions Architect will design efficient inference pipelines for dense and sparse Mixture-of-Experts (MoE) models, optimize performance across quantization, speculative decoding, and memory management, and collaborate with NVIDIA product teams to drive technical innovation. The position also includes leading technical workshops and shaping reference architectures to advance scalable AI inference in high-performance computing environments.