Senior Machine Learning Infrastructure Engineer, Research
This role involves building and operating scalable ML infrastructure for training and serving large physics models, with a focus on distributed training optimization, data pipeline performance, and model deployment. You'll work closely with research scientists and ML engineers in a high-impact R&D environment, solving systems-level challenges in GPU clusters and HPC workflows. The position emphasizes end-to-end ownership of research infrastructure, reproducibility, and developer experience for fast iteration.