Principal Machine Learning Infrastructure Engineer
This role involves designing and operating distributed machine learning infrastructure for large-scale physics-based AI models, with a focus on training efficiency, data pipeline performance, and model serving. The engineer will work closely with research scientists and ML engineers to build scalable systems on NVIDIA DGX hardware, optimize I/O for complex mesh datasets, and enable reliable deployment of models into customer environments. A strong foundation in distributed training, HPC systems, and Kubernetes is essential.