Research Engineer - Inference
The role involves deploying and optimizing cutting-edge AI models for real-time inference in production environments. You'll focus on improving latency, throughput, and cost efficiency across the full serving stack, from model architecture to custom kernels and orchestration. The position emphasizes building high-performance systems and tooling to enable rapid, reliable deployment of AI models at scale.