Senior Deep Learning Software Engineer, Inference
Design and optimize high-performance deep learning inference software for NVIDIA's GPU-accelerated platforms, focusing on frameworks like vLLM and SGLang. Improve model serving efficiency for large language and generative AI models across datacenter and edge devices. Collaborate with research and engineering teams to implement cutting-edge algorithms and performance optimizations using CUDA, Triton, and other low-level tools.