3/6/2025
What this post added
This post introduces Mistral OCR, an Optical Character Recognition API that comprehends documents including media, text, tables, and equations with high accuracy. It extracts content in an ordered interleaved text and images format, making it suitable for RAG systems with multimodal documents. The API is released as `mistral-ocr-latest` with pricing per page and a double page-per-dollar rate for batch inference. Key highlights include state-of-the-art understanding of complex documents, native multilingual and multimodal capabilities, top-tier benchmarks, fast processing speeds (up to 2000 pages per minute on a single node), 'doc-as-prompt' functionality for structured output (JSON), and a selective self-hosting option for sensitive data. Use cases include digitizing scientific research, preserving historical heritage, streamlining customer service, and making literature AI-ready.