
Extract text from documents and images with Datalab Marker and OCR
10/21/2025
Introduced Datalab Marker and OCR models to Replicate. Marker converts documents (PDF, DOCX, PPTX, images) to markdown or JSON, with features for formatting tables, math, code, and structured field extraction using JSON Schema. OCR detects text in 90 languages from images and documents, returning reading order and table grids. Both models are noted for their speed and accuracy, with Marker processing a page in ~0.18s and batching up to 120 pages/sec. Performance benchmarks on olmOCR-Bench are provided, showing Marker's superiority over other OCR solutions.