Document AI
H2OVL Mississippi

H2OVL Mississippi

6/24/2026

What this post added

Introduces H2OVL Mississippi-2B and 0.8B, new multimodal foundation models specifically for OCR and Document AI use cases. Details their architecture, inspired by LLaVA and InternVL, using a ViT-MLP-LLM setup with dynamic resolution and multi-scale adaptive cropping. Explains the two-stage training methodology, including pretraining on large datasets for image-text alignment and fine-tuning with specific tasks like QA, OCR, reasoning, and captioning. Highlights performance benchmarks showing superiority in text recognition and competitive performance in image benchmarks.

Read the original post ↗