Multimodal Safety Classification
[Deprecated] Pixtral 12B | Mistral AI

[Deprecated] Pixtral 12B | Mistral AI

9/17/2024

What this post added

This post announces Pixtral 12B, a natively multimodal model. It details its architecture, including a new 400M parameter vision encoder and a 12B parameter multimodal decoder based on Mistral Nemo. The model supports variable image sizes and aspect ratios, and can process multiple images within a 128k token context window. It demonstrates strong performance on multimodal reasoning benchmarks (MMMU), chart understanding, document question answering, and image-to-code generation, while also maintaining text-only benchmark performance. The post also notes that Pixtral 12B is deprecated and replaced by newer models.

Read the original post ↗