Serverless GPU Inference
Introducing: B200s and H200s on Modal

Introducing: B200s and H200s on Modal

5/30/2025

What this post added

This post announces the availability of NVIDIA B200 and H200 GPUs on the Modal platform. It details the specifications of these new GPUs, comparing them to H100s, and explains how their increased memory, memory bandwidth, and FP4 Tensor Core support benefit large model inference, particularly for Mixture-of-Experts (MoE) models. Benchmarks are provided showing significant latency and throughput improvements for LLM inference (DeepSeek V3 Large) on B200s compared to H200s. The post also highlights Modal's infrastructure advantages for GPU deployment, including fast spin-up times, autoscaling, and pay-as-you-go pricing.

Read the original post ↗