
6/26/2025 · Omar Sanseviero, Ian Ballantyne
What this post added
Introduced Gemma 3n, a mobile-first multimodal AI architecture optimized for on-device deployment. Key innovations include the MatFormer architecture for elastic inference and custom model sizing (Mix-n-Match), Per-Layer Embeddings (PLE) for memory efficiency by offloading embeddings to CPU, KV Cache Sharing for faster long-context processing, an advanced audio encoder based on Universal Speech Model (USM) enabling on-device ASR and AST, and the MobileNet-V5-300M vision encoder for high-performance multimodal tasks on edge devices. Gemma 3n models are available in E2B and E4B sizes, offering performance comparable to smaller traditional models with significantly reduced memory footprints.