Google Developer Platform
Introducing PaliGemma 2 mix: A vision-language model for multiple tasks- Google Developers Blog

Introducing PaliGemma 2 mix: A vision-language model for multiple tasks- Google Developers Blog

2/19/2025 · Omar Sanseviero, Andreas Steiner

What this post added

Introduced PaliGemma 2 mix, a vision-language model tuned for a mixture of tasks including image segmentation, captioning, OCR, object detection, and question answering. The model is available in multiple sizes (3B, 10B, 28B parameters) and resolutions (224px, 448px), and is compatible with Hugging Face Transformers, Keras, PyTorch, JAX, and Gemma.cpp. The post provides examples of using the model for detection, multiple object detection, OCR, segmentation, and question answering, demonstrating its out-of-the-box capabilities.

Read the original post ↗