Blogs›Mistral AI›Multimodal Safety Classification
Multimodal Safety Classification
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. It accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining, and delivers calibrated safety scores efficiently. Le Chat now integrates this capability, allowing for document and image analysis powered by the new Pixtral Large multimodal model, enabling summarization. Pixtral 12B is a natively multimodal model trained with interleaved image and text data, excelling in multimodal tasks and instruction following while maintaining state-of-the-art text-only performance. It features a new 400M parameter vision encoder and a 12B parameter multimodal decoder based on Mistral Nemo, supporting variable image sizes and multiple images within a 128k token context window. Pixtral 12B demonstrates strong performance on multimodal reasoning benchmarks like MMMU and excels in chart understanding, document question answering, and image-to-code generation.