Speculative Decoding Acceleration
AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators

AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators

10/10/2025

What this post added

Introduced ATLAS (AdapTive-LeArning Speculator System), a novel speculative decoding system that dynamically improves at runtime by learning from historical patterns and live traffic. ATLAS offers automatic performance improvements without manual tuning, continuously aligning with the target model's behaviors in real time. It achieves up to 2.65x speedup on DeepSeek-V3.1 and Kimi-K2, outperforming standard decoding and specialized hardware like Groq.

Read the original post ↗