Serverless GPU Inference
Modal + Mistral 3: 10x faster cold starts with GPU snapshotting

Modal + Mistral 3: 10x faster cold starts with GPU snapshotting

12/2/2025

What this post added

Introduced GPU memory snapshotting as an experimental feature to drastically reduce cold start times for GPU-intensive workloads. This feature, when enabled with sleep mode for vLLM servers, shifts GPU memory to CPU memory, allowing for faster restoration from snapshots. Tested on Ministral 3 3B, achieving a 10x reduction in median cold start time.

Read the original post ↗