Multimodal Safety Classification
Introducing Shieldstral. | Mistral AI

Introducing Shieldstral. | Mistral AI

8/4/2026

What this post added

Introduced Shieldstral, a 3B multimodal safety classifier. Developed a novel approach framing content moderation as a policy-adaptive question-answering task, allowing for plain-language policies at inference time. Unified text and image safety evaluation without retraining. Achieved strong performance on text safety, refusal detection, policy adaptability, and multimodal benchmarks. Built using Forge, Mistral AI's platform for training, aligning, and evaluating custom models.

Read the original post ↗