Senior Deep Learning Software Engineer, Inference
Design and optimize GPU-accelerated deep learning inference software for large-scale language and generative AI models. Contribute to open-source frameworks like vLLM and SGLang, and implement performance improvements across NVIDIA's full range of accelerators from datacenter to edge. Work closely with research and engineering teams to profile, tune, and deploy state-of-the-art model serving solutions using CUDA, Triton, and other low-level optimization tools.