8/4/2026
What this post added
Introduced Shieldstral, a 3B multimodal safety classifier. Developed a novel approach framing content moderation as a policy-adaptive question-answering task, allowing for plain-language policies at inference time. Unified text and image safety evaluation without retraining. Achieved strong performance on text safety, refusal detection, policy adaptability, and multimodal benchmarks. Built using Forge, Mistral AI's platform for training, aligning, and evaluating custom models.