
6/3/2026 · André Susano Pinto, Andreas Steiner, Karolis Misiunas, Karsten Roth, Michael Tschannen, Omar Sanseviero
What this post added
Introduced Gemma 4 12B, a dense multimodal model with a unified, encoder-free architecture. This model features reduced multimodal latency by bypassing separate vision and audio encoders, is the first medium-sized Gemma model with native audio input, and is designed for local inference on GPUs with 16GB VRAM. A new macOS desktop experience is also released. The architecture utilizes a single decoder-only transformer with a vision embedder (35M parameters) and linear projection for audio wave input. Unified fine-tuning allows for updating the entire multimodal token loop. The model demonstrates capabilities in automatic speech recognition, agentic reasoning, diarization, video understanding, and coding. On-device and desktop serving are powered by LiteRT-LM, including native macOS apps and OpenAI-compatible local API servers. Getting started resources include LM Studio, Ollama, Google AI Edge Gallery App, Google AI Edge Eloquent app, LiteRT-LM CLI, Hugging Face, Kaggle, developer documentation, quick start notebooks, and integrations with Transformers, llama.cpp, MLX, SGLang, vLLM, and Unsloth. A Gemma Skills Repository is released to support agentic development.