Senior Solutions Architect – Large Scale AI Inference
This role involves guiding EMEA-based AI companies in deploying and optimizing large-scale AI inference workloads on multi-node GPU clusters. The Senior Solutions Architect will design efficient inference pipelines for dense and sparse Mixture-of-Experts (MoE) models, improve performance through quantization and speculative decoding, and collaborate with NVIDIA’s product teams. The position also includes leading technical workshops and shaping next-generation inference architectures at the intersection of AI and high-performance computing.