Senior Software Engineer, AI Inference Systems
Design and optimize high-performance AI inference systems for large-scale models using cutting-edge GPU hardware. Develop and enhance inference frameworks like vLLM, implement advanced parallelism techniques, and contribute to compiler and kernel optimization. Build scalable scheduling solutions for multi-node, multi-cloud GPU deployments and drive industry benchmarks such as MLPerf.