
3/26/2026 · Cheney Zhang
What this post added
Introduces the CCKM benchmark (Cross-modal, Cross-lingual, Key information, MRL dimension compression) for evaluating embedding models in production RAG scenarios. Benchmarks 10 embedding models across cross-modal retrieval, cross-lingual retrieval, and key information retrieval, and analyzes dimension compression impact. Identifies Gemini Embedding 2 as a strong all-rounder, Qwen3-VL-2B for cross-modal tasks, and Voyage Multimodal 3.5/Jina Embeddings v4 for dimension compression. Discusses modality gap and semantic understanding as key factors for model performance in RAG.