Senior Deep Learning Software Engineer, Inference
This role involves designing, optimizing, and maintaining high-performance deep learning inference software for large-scale language and generative AI models. The engineer will work on GPU-accelerated frameworks like vLLM, SGLang, and FlashInfer, implementing cutting-edge algorithms and performance improvements across NVIDIA's full range of accelerators. Collaboration with open-source communities and cross-team partners is central to advancing model serving efficiency and deployment at scale.