
4/30/2025 · Ju-yeong Ji, Ravin Kumar
What this post added
This post introduces Gemma 3, detailing its new vision-language capabilities enabled by a SigLIP encoder and a "Pan&Scan" algorithm for image processing. It explains the use of "soft tokens" to reduce inference resource requirements. Architectural improvements include 5-to-1 interleaved attention for better context handling and QK-norm for improved accuracy and speed. The post highlights Gemma 3's extended context length support (up to 128k tokens) and its use of bidirectional attention for image inputs, contrasting it with PaliGemma's autoregressive approach. A new tokenizer with a 262k vocabulary size is also introduced for enhanced multilingual support.