
6/23/2026
What this post added
Introduces DFlash, an open-source block diffusion model for speculative decoding that accelerates LLM inference by drafting entire token blocks in parallel. Demonstrates up to 15x throughput improvement on NVIDIA Blackwell GPUs for gpt-oss-120b and nearly doubles interactivity for Llama 3.1 8B compared to EAGLE-3. Highlights integration with SGLang, vLLM, and TensorRT-LLM, enabling adoption without code refactoring.