Senior Solutions Architect – Large Scale Neural Networks Inference
Lead technical strategy for AI inference deployments across EMEA, working closely with AI-native enterprises and frontier labs to optimize large-scale neural network inference on NVIDIA's platform. Architect high-performance inference pipelines using frameworks like TensorRT-LLM, vLLM, and SGLang, and translate real-world deployment challenges into product improvements. Focus on solving critical issues around latency, efficiency, memory utilization, and GPU cluster performance.