Staff Software Engineer, Inference
This role involves designing and maintaining large-scale, performance-sensitive distributed systems that serve Claude to millions of users globally. The engineer will work across the full inference stack, building intelligent routing, autoscaling, and deployment systems for AI models running on diverse accelerators across multiple cloud platforms. The position emphasizes real-time system resilience, compute efficiency, and close collaboration with research teams to enable next-generation model development.