3/16/2026
What this post added
Mistral Small 4 introduces a unified model architecture combining reasoning, multimodal, and agentic coding capabilities. It utilizes a Mixture of Experts (MoE) design with 128 experts, 4 active per token, and has 119B total parameters with 6B active. The model supports a 256k context window and features a configurable `reasoning_effort` parameter. Performance improvements include a 40% reduction in latency and a 3x increase in throughput compared to its predecessor. Native multimodality is integrated, accepting both text and image inputs. The post details hardware requirements for enterprise deployment and highlights optimizations for vLLM and SGLang through collaboration with NVIDIA.