Senior Software Engineer, RL Post-Training Frameworks
This role involves designing and building scalable reinforcement learning post-training infrastructure that supports the full lifecycle of RL workflows, from experimentation to production at scale. The engineer will optimize distributed training-inference-rollout loops across heterogeneous hardware, contribute to open-source RL frameworks like VeRL and TorchTitan, and collaborate with research and hardware teams to shape future AI capabilities. Key challenges include fault tolerance, elastic scaling, and efficient coordination of actor-critic-reward models in complex distributed environments.