Senior HPC AI Cluster Engineer
This role involves designing, implementing, and maintaining large-scale HPC and AI clusters, with a focus on automation, monitoring, and performance optimization across the full stack—from bare metal to application level. The engineer will develop CI/CD pipelines, manage orchestration tools like Slurm and Kubernetes, and collaborate with researchers and developers to improve workflows on cutting-edge GPU-accelerated platforms. A strong foundation in Linux, networking, storage, and scripting is essential, along with experience in deploying scalable infrastructure solutions.