
6/29/2026
What this post added
Introduces Gemma 4 31B, a multimodal model, to the Cerebras Inference platform, achieving over 1,800 tokens per second. This represents a significant advancement in multimodal AI inference speed and latency, enabling real-time visual and agentic workflows. The post highlights the model's capabilities in image understanding combined with wafer-scale speed, unlocking new product experiences like screenshot-to-insight, long-context summarization, and screenshot-to-patch generation. It also positions Gemma 4 as a reference medium-size model on Cerebras, comparable in intelligence to Claude Haiku but significantly faster.