Senior HPC Performance Engineer
Conduct performance characterization and analysis on large multi-GPU and multi-node clusters, focusing on GPU communication libraries like NCCL and NVSHMEM. Evaluate system-level interactions across hardware and software stacks, develop tools to visualize performance data, and debug issues across the full stack. Work closely with cross-functional teams to optimize communication performance at scale for HPC and deep learning applications.