BlogsReplicateDocument Parsing and OCR

Document Parsing and OCR

Document Parsing and OCR

1
posts
2025

Replicate now offers advanced document parsing and text extraction capabilities through the integration of Datalab Marker and OCR models. Marker processes various document formats (PDF, DOCX, PPTX, images) into markdown or JSON, handling tables, math, and code, and supporting structured field extraction via JSON Schema. OCR detects text in 90 languages from images and documents, providing reading order and table grid data. Both models offer high performance, outperforming established tools like Tesseract, with Marker achieving up to 120 pages per second when batched. Performance benchmarks show Marker (Balanced mode) achieving an overall score of 82.7 ± 0.9 on the olmOCR-Bench, surpassing other models including GPT-4o and Deepseek OCR.

2025

Extract text from documents and images with Datalab Marker and OCR

10/21/2025

Introduced Datalab Marker and OCR models to Replicate. Marker converts documents (PDF, DOCX, PPTX, images) to markdown or JSON, with features for formatting tables, math, code, and structured field extraction using JSON Schema. OCR detects text in 90 languages from images and documents, returning reading order and table grids. Both models are noted for their speed and accuracy, with Marker processing a page in ~0.18s and batching up to 120 pages/sec. Performance benchmarks on olmOCR-Bench are provided, showing Marker's superiority over other OCR solutions.