BlogsMistral AIVoxtral Speech Understanding Models

Voxtral Speech Understanding Models

Voxtral Speech Understanding Models

3
posts
2025–2026

Shieldstral introduces Voxtral Transcribe 2, featuring Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live applications. Voxtral Realtime offers ultra-low latency (sub-200ms) with a novel streaming architecture and is open-weights under Apache 2.0. Voxtral Mini Transcribe V2 provides state-of-the-art transcription with speaker diarization, context biasing, and word-level timestamps in 13 languages, achieving industry-leading accuracy at a low cost. Both models demonstrate strong multilingual performance and noise robustness. An audio playground in Mistral Studio allows for instant testing.

2026

Speaking of Voxtral | Mistral AI

3/23/2026

This post introduces Voxtral TTS, a new text-to-speech model. It details the model's architecture (transformer-based, autoregressive, flow-matching, built on Ministral 3B, with specific parameter counts for backbone, acoustic transformer, and codec), its performance metrics (low latency, RTF, multilingual capabilities, voice adaptation), and its use cases in enterprise voice workflows. It also includes comparative human evaluation data against ElevenLabs.

Voxtral transcribes at the speed of sound. | Mistral AI

2/4/2026

Introduces Voxtral Transcribe 2, comprising Voxtral Mini Transcribe V2 (batch) and Voxtral Realtime (streaming). Voxtral Realtime features a novel streaming architecture for sub-200ms latency and is released under Apache 2.0. Voxtral Mini Transcribe V2 enhances transcription quality with speaker diarization, context biasing, and word-level timestamps across 13 languages. Both models show improved performance and efficiency compared to previous versions and competitors.

2025

Voxtral | Mistral AI

7/15/2025

This post introduces the Voxtral models, detailing their architecture (24B and 3B variants), performance benchmarks against leading models (Whisper, GPT-4o mini, Gemini 2.5 Flash, ElevenLabs Scribe), and key features like long-form context, multilingual support, and function calling. It highlights the technical trade-offs addressed by Voxtral in the speech intelligence market and outlines enterprise-grade features and future development plans for audio capabilities.