BlogsMistral AIMistral OCR

Mistral OCR

Mistral OCR

3
posts
2025–2026

Mistral OCR 4 introduces breakthrough document parsing with bounding boxes, block classification, and inline confidence scores, supporting 170 languages and running in a single container for self-hosted deployments. It outperforms leading OCR systems in human evaluations and benchmarks while offering cost-efficient, high-throughput processing for enterprise search, RAG, and agentic workflows. Available via API or Document AI, it provides structured outputs for custom pipelines or no-code applications, with self-hosting options for data privacy.

2026

Mistral OCR 4 : SOTA OCR for Document Intelligence

6/23/2026

Mistral OCR 4 introduces bounding boxes, block classification, and inline confidence scores, expanding beyond text extraction. It supports 170 languages, can be self-hosted in a single container, and integrates with the Mistral Search Toolkit. Performance improvements are highlighted through human preference evaluations and benchmarks like OlmOCRBench and OmniDocBench, with detailed breakdowns for multilingual performance. The post also discusses recommended use cases and limitations.

2025

Introducing Mistral OCR 3 | Mistral AI

12/17/2025

This post introduces Mistral OCR 3, detailing its performance improvements over Mistral OCR 2, including a 74% overall win rate. It highlights the model's state-of-the-art accuracy, its ability to handle diverse document types (forms, handwriting, complex tables, low-quality scans), and its output format (markdown with HTML table reconstruction). The post also mentions its competitive pricing and integration into Mistral AI Studio's Document AI Playground. Benchmarks and specific upgrades in handwriting, form, scanned document, and table processing are detailed.

Mistral OCR | Mistral AI

3/6/2025

This post introduces Mistral OCR, an Optical Character Recognition API that comprehends documents including media, text, tables, and equations with high accuracy. It extracts content in an ordered interleaved text and images format, making it suitable for RAG systems with multimodal documents. The API is released as `mistral-ocr-latest` with pricing per page and a double page-per-dollar rate for batch inference. Key highlights include state-of-the-art understanding of complex documents, native multilingual and multimodal capabilities, top-tier benchmarks, fast processing speeds (up to 2000 pages per minute on a single node), 'doc-as-prompt' functionality for structured output (JSON), and a selective self-hosting option for sensitive data. Use cases include digitizing scientific research, preserving historical heritage, streamlining customer service, and making literature AI-ready.