Senior Solutions Architect – Large Scale AI Inference
This role involves guiding AI-native customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters, with a focus on architecting efficient pipelines for dense and sparse Mixture-of-Experts models. The Senior Solutions Architect will tackle challenges in quantization, speculative decoding, KV cache management, and interconnect-aware scheduling, while collaborating closely with NVIDIA product teams and leading technical communities across EMEA. Work centers on advancing high-performance AI inference systems at scale, particularly in memory-bound and distributed environments.