Senior Software Engineer, AI Inference Systems
Design and optimize high-performance AI inference systems for large-scale models on NVIDIA GPUs. Develop and enhance inference frameworks like vLLM, optimize GPU kernels and compilers, and build scalable scheduling solutions across multi-node, multi-cloud environments. Contribute to benchmarking standards such as MLPerf and integrate cutting-edge research into production software to advance the state of accelerated computing.