7/16/2026
What this post added
This post details significant performance optimizations for real-time video generation inference on Baseten, specifically for the Wan 2.2 model. Key contributions include timestep distillation to reduce generation steps, custom kernel engineering for optimized attention mechanisms (e.g., Video Sparse Attention), and NVFP4 quantization for improved memory bandwidth and tensor core throughput. The post also highlights the scalable inference infrastructure required to support this capability, including optimized cold starts via the Baseten Delivery Network and intelligent queuing, as well as the implementation of custom content guardrails for responsible deployment.