Software Engineer (Compute Efficiency), London
This role involves designing, deploying, and scaling robust observability systems and telemetry pipelines to monitor compute efficiency and hardware health across distributed clusters. You will work closely with ML and platform teams to optimize accelerator utilization and improve workload goodput, contributing to the reliability of ML runs and the overall efficiency of the compute infrastructure.