
3/20/2026
What this post added
This post highlights the increasing importance of inference speed in the AI race, driven by its impact on model iteration and software development velocity. It provides examples of major AI labs prioritizing speed, such as Google's Gemini 3 Flash, Anthropic's faster Claude Opus 4.6, and OpenAI's partnership with Cerebras for GPT-5.3-Codex-Spark. The post explains how fast inference enables recursive self-improvement of AI models and accelerates product development cycles, citing Anthropic's pricing strategy and a hypothetical scenario of two companies developing an AI-powered CRM. It connects this trend to the broader digital economy's historical focus on speed and positions high-speed inference as critical infrastructure for the AI era.