On-device Audio Synthesis
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

7/28/2026

What this post added

Introduces a novel memory-efficient audio synthesis architecture for on-device deployment, featuring a decoupled temporal and depth processing design using RVQ. Key innovations include a single reusable depth decoder with DiT-style conditioning and causal sliding window attention with fixed-window key-value caching for constant memory complexity. This architecture enables real-time synthesis within tight resource constraints and has been deployed in Siri Expressive Voices, demonstrating significant improvements in audio quality metrics.

Read the original post ↗