Predictive APIs and Machine Learning Integration
Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings- Google Developers Blog

Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings- Google Developers Blog

9/4/2025 · Min Choi, Sahil Dua, Alice Lisak

What this post added

Introduces EmbeddingGemma, a new open embedding model based on Gemma 3 architecture, optimized for on-device AI. It delivers best-in-class performance for its size (308M parameters), supporting over 100 languages and enabling RAG and semantic search on-device. Key features include flexible output dimensions via Matryoshka Representation Learning, low RAM usage (<200MB) with quantization, and fast inference times (<15ms on EdgeTPU). It integrates with popular tools like sentence-transformers, llama.cpp, and LangChain.

Read the original post ↗