Inference Performance Optimization
Runway chooses Modal to power real-time inference for Runway Characters | Modal Blog

Runway chooses Modal to power real-time inference for Runway Characters | Modal Blog

3/26/2026

What this post added

This post details how Runway leverages Modal's platform for real-time inference of Runway Characters, a video agent API. It highlights Modal's ability to handle GPU-intensive, latency-critical, and variable demand workloads. Specifically, it mentions Runway's ability to turn containers into multi-node GPU clusters with RDMA networking via a single line of code, enabling distribution of inference across multiple GPUs with high-bandwidth communication between nodes. This allows for sustained low latency across the full duration of a conversation, with expressions, lip-sync, and gestures, without degradation, and deployment across global regions at production scale.

Read the original post ↗