
2/19/2025 · Omar Sanseviero, Andreas Steiner
What this post added
Introduced PaliGemma 2 mix, a vision-language model tuned for a mixture of tasks including image segmentation, captioning, OCR, object detection, and question answering. The model is available in multiple sizes (3B, 10B, 28B parameters) and resolutions (224px, 448px), and is compatible with Hugging Face Transformers, Keras, PyTorch, JAX, and Gemma.cpp. The post provides examples of using the model for detection, multiple object detection, OCR, segmentation, and question answering, demonstrating its out-of-the-box capabilities.