Senior Software Engineer, AI Inference Systems
This role involves building and optimizing high-performance AI inference systems for large-scale models, focusing on GPU kernel optimization, compiler development, and distributed inference frameworks. You'll work on cutting-edge technologies like vLLM, speculative decoding, and multi-node GPU orchestration, while contributing to industry benchmarks like MLPerf. The position blends deep systems engineering with research to advance the efficiency and scalability of AI inference.