BlogsMistral AI Feature Trails

Mistral AI logo

Mistral AI Feature Trails

See how major capabilities shipped, upgraded, and evolved across Mistral AI's engineering blog.

Feature trails

21

Multimodal Safety Classification

Active

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. It accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining, and delivers calibrated safety scores efficiently. Le Chat now integrates this capability, allowing for document and image analysis powered by the new Pixtral Large multimodal model, enabling summarization. Pixtral 12B is a natively multimodal model trained with interleaved image and text data, excelling in multimodal tasks and instruction following while maintaining state-of-the-art text-only performance. It features a new 400M parameter vision encoder and a 12B parameter multimodal decoder based on Mistral Nemo, supporting variable image sizes and multiple images within a 128k token context window. Pixtral 12B demonstrates strong performance on multimodal reasoning benchmarks like MMMU and excels in chart understanding, document question answering, and image-to-code generation.

5 posts

Timeline

20242026

Mistral Agents API

Active

Mistral AI Studio is introduced as a comprehensive production AI platform, building upon the Mistral Agents API. It integrates Observability, Agent Runtime, and AI Registry to provide a unified system for building, evaluating, and running AI in production. Key features include detailed traffic inspection, dataset building, evaluation logic definition, traceable feedback loops, durable and reproducible agent execution via Temporal, and a system of record for all AI assets with lineage tracking. This post adds version control for prompts and skills within Studio, treating them as versioned, owned, and traceable production assets with immutable versions, rollback capabilities, clear ownership, classification labels, and audit logs. This enables faster iteration and controlled deployment of AI behavior.

5 posts

Timeline

20252026

Robostral Navigate Embodied Navigation

Active

Mistral AI introduces Robostral Navigate, an 8B model for embodied navigation using a single RGB camera. It achieves state-of-the-art performance on R2R-CE benchmarks, outperforming multi-sensor approaches. The model is trained in-house using simulated data and token-efficient techniques like prefix-caching, and further improved with online reinforcement learning (CISPO). It generalizes across robot types and adapts to real-world obstacles, combining pointing-based navigation with local displacement fallbacks.

1 post

Timeline

Formal Verification Models

Active

Mistral AI introduces Leanstral 1.5, a 6B active parameter, Apache-2.0 licensed model for formal verification. It achieves state-of-the-art results on benchmarks like miniF2F, PutnamBench, FATE-H, and FATE-X. The model is trained using a three-stage process including mid-training, supervised fine-tuning, and reinforcement learning with CISPO, enabling it to perform agentic proof engineering and real-world code verification. Leanstral 1.5 has demonstrated the ability to uncover previously unknown bugs in open-source repositories.

1 post

Timeline

Connectors and Tool Orchestration

Active

Mistral AI introduces enhanced controls for Connectors, including enriched admin controls for workspace/org-level access and tool-level permissions, API keys with connector scopes for secure automated workloads, multi-account connectors for authenticating with multiple accounts, a Connectors Debugger for root cause analysis of MCP connectors, and integration of Connectors into Vibe Code and Workflows for seamless and governed access to enterprise tools. This expands on the initial release of Connectors and tool orchestration by adding granular security and debugging capabilities.

2 posts

Timeline

Mistral OCR

Active

Mistral OCR 4 introduces breakthrough document parsing with bounding boxes, block classification, and inline confidence scores, supporting 170 languages and running in a single container for self-hosted deployments. It outperforms leading OCR systems in human evaluations and benchmarks while offering cost-efficient, high-throughput processing for enterprise search, RAG, and agentic workflows. Available via API or Document AI, it provides structured outputs for custom pipelines or no-code applications, with self-hosting options for data privacy.

3 posts

Timeline

20252026

Mistral Vibe Terminal Agent

Active

Mistral Vibe now supports remote agents running in the cloud, powered by Mistral Medium 3.5. This enables parallel, asynchronous task execution for coding and productivity work. A new 'Work mode' in Le Chat allows for complex, multi-step tasks involving cross-tool workflows, research, and inbox triage, with human oversight for sensitive actions. Mistral Medium 3.5 is a 128B dense model with a 256k context window, optimized for long-horizon tasks, instruction-following, reasoning, and coding. Le Chat also introduces Deep Research mode for structured reports, Voice mode powered by Voxtral, Projects for organizing conversations, and advanced image editing capabilities.

6 posts

Timeline

20252026

Search Toolkit

Active

Mistral AI introduces Search Toolkit, a composable, open-source framework for building production search pipelines for AI applications. It integrates ingestion, retrieval (BM25, dense embedding, hybrid), and evaluation (recall, precision, MRR, NDCG) into a single framework with a shared interface, aiming to reduce engineering time spent on plumbing and integrations. It supports various deployment environments (cloud, on-premises, edge) and is designed for enterprise search, RAG, and domain-specific applications. This post details using LLM-as-a-judge with the RAG Triad framework and Mistral's structured outputs to evaluate RAG systems, enhancing the evaluation capabilities within the Search Toolkit.

2 posts

Timeline

20252026

Physics AI for Engineering

Active

Mistral AI introduces physics AI, a new foundational capability for AI-native industrial engineering. This capability leverages data-driven AI models to predict physical behavior directly from geometry and boundary conditions, offering a significant acceleration over traditional numerical physics simulations. The technology enables faster product design by exploring thousands of design variants, accelerated tooling and process design by optimizing tooling geometry and parameters, and real-time d. With the acquisition of Emmi AI, Mistral significantly enriches its products and expertise in this domain, aiming to accelerate the work of engineering solution teams worldwide. The acquisition also accelerates our Science roadmap, advancing our understanding of fundamental physics and leveraging unique industrial data. With Emmi AI's models complementing our own, we are set to build the best-in-class agents for engineers.

3 posts

Timeline

Voxtral Speech Understanding Models

Active

Shieldstral introduces Voxtral Transcribe 2, featuring Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live applications. Voxtral Realtime offers ultra-low latency (sub-200ms) with a novel streaming architecture and is open-weights under Apache 2.0. Voxtral Mini Transcribe V2 provides state-of-the-art transcription with speaker diarization, context biasing, and word-level timestamps in 13 languages, achieving industry-leading accuracy at a low cost. Both models demonstrate strong multilingual performance and noise robustness. An audio playground in Mistral Studio allows for instant testing.

3 posts

Timeline

20252026

Model Customization and Agents

Active

Mistral AI introduces Mistral Saba, a specialized 24B parameter regional language model trained on datasets from the Middle East and South Asia, excelling in Arabic and South Indian languages. It offers superior accuracy and cost-efficiency compared to larger models and can be deployed via API or locally on single-GPU systems. This expands the model customization offerings by providing a pre-trained regional model that can serve as a base for further fine-tuning for domain-specific expertise and culturally relevant content creation.

11 posts

Timeline

20232026

Mixtral Sparse Mixture of Experts Model

Active

Mistral AI introduces Mixtral 8x22B, a new open-source sparse Mixture-of-Experts (SMoE) model. It utilizes 39B active parameters out of 141B, offering significant cost efficiency and performance. The model is fluent in English, French, Italian, German, and Spanish, with strong mathematics and coding capabilities. It natively supports function calling and has a 64K token context window. Released under Apache 2.0, it aims to provide unmatched cost efficiency and performance for its size, serving as a strong base for fine-tuning.

11 posts

Timeline

20232026

vLLM Memory Leak Debugging

Active

This post details the debugging of a memory leak in vLLM, specifically within a disaggregated Prefill/Decode serving setup utilizing NIXL and UCX. The investigation involved advanced profiling tools like Heaptrack and kernel-level tracing with pmap and BPFtrace to identify memory growth outside of the traditional heap, specifically in anonymous memory mappings managed by mmap/mremap. The root cause was identified as a leak within the KV Cache transfer mechanism via NIXL, leading to unreleased memory regions.

1 post

Timeline

AI Memory System

Active

Mistral AI introduces a new AI memory system for Le Chat, focusing on transparency, user agency, and data sovereignty. The system allows users to control what the AI remembers, how it recalls information, and provides visibility into the sources of recalled data. It features automatic saving of useful information with smart, timely, and visible recall, user controls for turning memory on/off, incognito modes, and editing/deleting memories. The system is designed to be portable and interoperable. This update enhances the memory system with user controls for editing/deleting memories and import capabilities from ChatGPT.

2 posts

Timeline

20252026

Codestral Mamba Architecture

Cooling

Mistral AI introduces Codestral 25.08, an updated version of its code generation model, featuring a 30% increase in accepted completions, 10% more retained code, and 50% fewer runaway generations. It also includes improvements in chat mode for instruction following and code abilities. The post also details the complete Mistral Coding Stack for Enterprise, integrating Codestral Embed for semantic retrieval, Devstral for autonomous multi-step development, and Mistral Code for IDE integration. The stack emphasizes self-hosting, privacy, and observability for enterprise adoption.

5 posts

Timeline

20242026

Mistral Compute AI Infrastructure

Cooling

Mistral AI launches Mistral Compute, a new AI infrastructure offering providing customers with a private, integrated stack of GPUs, orchestration, APIs, products, and services. This offering aims to democratize AI infrastructure by providing an alternative to existing cloud providers, with a focus on sovereign AI capabilities, sustainability, and data sovereignty. It includes access to NVIDIA reference architectures and tens of thousands of GPUs, designed to support training and serving of AI workloads.

1 post

Timeline

20252026

Magistral Reasoning Model

Cooling

Mistral AI introduces Magistral, a new family of reasoning models designed for domain-specific, transparent, and multilingual reasoning. The models are released in two variants: Magistral Small (24B parameters, open-source) and Magistral Medium (enterprise version). Magistral excels in multi-step logic, providing traceable thought processes and high-fidelity reasoning across multiple languages. It is optimized for speed, with Magistral Medium achieving up to 10x faster token throughput in Le Chat. This post introduces MathΣtral, a specialized 7B parameter model derived from Mistral 7B, focusing on advanced mathematical problems requiring complex, multi-step logical reasoning. MathΣtral achieves state-of-the-art reasoning capacities in its size category across industry-standard benchmarks like MATH and MMLU, demonstrating significant performance improvements over its base model in STEM subjects. It can achieve even better results with increased inference-time computation, such as majority voting or using a strong reward model.

2 posts

Timeline

20242026

Codestral Embeddings

Cooling

Mistral AI releases Codestral Embed, a specialized embedding model for code. It offers superior performance for retrieval use cases on real-world code data compared to leading competitors like Voyage Code 3, Cohere Embed v4.0, and OpenAI's large embedding model. The model supports outputting embeddings with different dimensions and precisions, allowing for trade-offs between retrieval quality and storage costs. Codestral Embed is optimized for high-performance code retrieval and semantic understanding, enabling applications such as retrieval-augmented generation for code completion and editing, semantic code search, similarity search and duplicate detection, and semantic clustering and code analytics. It is available via API and for on-prem deployments, with recommendations for chunking strategies for retrieval use cases.

1 post

Timeline

20252026

Batch API for Model Inference

Cooling

Mistral AI introduces a Batch API for its models, offering a more cost-efficient way (50% lower cost) to process high-volume requests compared to synchronous API calls. This feature is ideal for applications prioritizing data volume over synchronous responses, enabling users to upload batch files and download processed outputs. Popular use cases include bulk customer feedback analysis, document summarization and translation, vector embedding generation, and data labeling. The Batch API is available for all models on La Plateforme and will be extended to cloud provider partners. Usage is limited to 1 million ongoing requests per workspace.

1 post

Timeline

20242026

Ministral Edge Models

Cooling

Mistral AI introduces Ministral 3B and Ministral 8B, new state-of-the-art models optimized for on-device and edge computing. These models offer enhanced knowledge, commonsense reasoning, function-calling, and efficiency in the sub-10B parameter category. They support up to 128k context length, with Ministral 8B featuring an interleaved sliding-window attention pattern for improved inference speed and memory efficiency. Use cases include privacy-first local inference for applications like smart assistants, local analytics, and autonomous robotics, as well as acting as efficient intermediaries for function-calling in multi-step agentic workflows. Benchmarks show these models consistently outperform peers in their size category, including Gemma 2 and Llama 3 variants, and even surpass Mistral 7B on many metrics. The models are available via API on la Plateforme with competitive pricing and also offered under Mistral Commercial and Research Licenses for self-deployment, with support for lossless quantization.

1 post

Timeline

20242026

Mistral Large Model

Quiet since 2024

Mistral AI releases Mistral Large, a new flagship text generation model with top-tier reasoning capabilities for complex multilingual tasks, text understanding, transformation, and code generation. It is natively fluent in English, French, Spanish, German, and Italian, features a 32K token context window, precise instruction-following, and native function calling. Mistral Large is available through la Plateforme and Azure. A new optimized model, Mistral Small, is also released for low-latency workloads, outperforming Mixtral 8x7B and offering similar RAG-enablement and function calling capabilities. The platform now offers simplified endpoint offerings with open-weight and optimized model endpoints, along with improvements in multi-currency pricing, service tiers, and reduced latency across all endpoints. JSON format mode and function calling are now available on mistral-small and mistral-large, with plans to extend to all endpoints.

1 post

Timeline

20242026