Baseten AI Model Deployment and Serving Platform
Real-time video generation inference on Baseten

Real-time video generation inference on Baseten

7/16/2026

What this post added

This post details significant performance optimizations for real-time video generation inference on Baseten, specifically for the Wan 2.2 model. Key contributions include timestep distillation to reduce generation steps, custom kernel engineering for optimized attention mechanisms (e.g., Video Sparse Attention), and NVFP4 quantization for improved memory bandwidth and tensor core throughput. The post also highlights the scalable inference infrastructure required to support this capability, including optimized cold starts via the Baseten Delivery Network and intelligent queuing, as well as the implementation of custom content guardrails for responsible deployment.

Read the original post ↗