Serverless GPU Inference
How Reducto improved enterprise-scale document processing latency by 3x

How Reducto improved enterprise-scale document processing latency by 3x

11/19/2025

What this post added

Reducto leveraged Modal's GPU memory snapshotting feature to reduce cold boot times for their inference models by 83%, from approximately 70 seconds to 12 seconds. This significantly improved their P90 latency by 3x for enterprise-scale document processing.

Read the original post ↗