
3/12/2025 · Omar Sanseviero, Philipp Schmid
What this post added
Introduces Gemma 3, the latest iteration of Google's open-model family, significantly enhancing capabilities with multimodality (vision-language input, text output), extended context windows (up to 128k tokens), broader language support (140+ languages), and improved math, reasoning, and chat functionalities. Details the optimized pre-training and post-training processes for Gemma 3, including distillation, RLHF, RLMF, and RLEF, and the use of a new tokenizer and large token counts on TPUs with JAX. Highlights the integrated vision encoder based on SigLIP for image and video analysis and the introduction of ShieldGemma 2 for image safety classification. Provides examples of multi-turn text and interleaved image inputs, and lists various tools and deployment options for Gemma 3.