
5/30/2025
What this post added
This post announces the availability of NVIDIA B200 and H200 GPUs on the Modal platform. It details the specifications of these new GPUs, comparing them to H100s, and explains how their increased memory, memory bandwidth, and FP4 Tensor Core support benefit large model inference, particularly for Mixture-of-Experts (MoE) models. Benchmarks are provided showing significant latency and throughput improvements for LLM inference (DeepSeek V3 Large) on B200s compared to H200s. The post also highlights Modal's infrastructure advantages for GPU deployment, including fast spin-up times, autoscaling, and pay-as-you-go pricing.