Senior Software Engineer, AI Inference Systems
Design and optimize high-performance AI inference systems for large-scale models using NVIDIA's latest GPU hardware. Develop and enhance inference frameworks like vLLM, implement advanced parallelism techniques, and build optimized GPU kernels. Contribute to compiler infrastructure, benchmarking, and distributed scheduling for multi-node, multi-cloud deployments. Integrate research innovations into production software and publish original work in ML systems.