.webp)
5/20/2026
What this post added
This post details a six-week benchmarking study on NVIDIA H100 GPUs at CoreWeave, focusing on achieving high Model FLOPs Utilization (MFU) and Mean Time To Failure (MTTF) for large-scale AI training. It highlights specific infrastructure optimizations including a dual-fabric network (InfiniBand + DPU-offloaded Ethernet), topology-aware scheduling with SUNK, automated re-queuing of failed jobs, and Tensorizer-based asynchronous checkpointing which reduced save times from 129s to 17s. The study achieved 51-52% MFU (vs. 35-45% typical) and 3.66 days MTTF at 1,024 GPUs (10x improvement), with projected improvements at 16,384 GPUs. It also details custom tokenizer performance and third-party validation against published results.