AI Inference Latency Optimization
Cerebras

Cerebras

12/4/2025

What this post added

This post details several research papers presented at NeurIPS 2025 that advance AI inference and training. Key contributions include: CODA (Conductor-Driven Architecture) for optimizing test-time compute by separating planning and execution roles, leading to smaller models outperforming larger ones; Calibrated Reasoning, an Explanatory Verifier trained via GRPO that provides calibrated confidence scores and natural language reasoning for solution correctness, enabling token savings in best-of-n strategies; DREAM, a speculative decoding framework for vision-language models that uses cross-attention and entropy-adaptive feature selection for speedup; CompleteP, a parameterization that enables hyperparameter transfer across model depths and widths, leading to compute savings; Power Lines, which identifies scaling laws for weight decay and batch size in LLM pre-training, enabling prediction of optimal settings; and PTPP-Aware Adaptation Scaling Laws that predict domain-adaptation performance based on pre-training budgets.

Read the original post ↗