ColPali: Efficient Document Retrieval with Vision Language Models
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo
摘要
Documents are visually rich structures that convey information through text, but also figures, page layouts, tables, or even fonts. Since modern retrieval systems mainly rely on the textual information they extract from document pages to index documents -often through lengthy and brittle processes-, they struggle to exploit key visual cues efficiently. This limits their capabilities in many practical document retrieval applications such as Retrieval Augmented Generation (RAG). To benchmark current systems on visually rich document retrieval, we introduce the Visual Document Retrieval Benchmark ViDoRe, composed of various page-level retrieval tasks spanning multiple domains, languages, and practical settings. The inherent complexity and performance shortcomings of modern systems motivate a new concept; doing document retrieval by directly embedding the images of the document pages. We release ColPali, a Vision Language Model trained to produce high-quality multi-vector embeddings from images of document pages. Combined with a late interaction matching mechanism, ColPali largely outperforms modern document retrieval pipelines while being drastically simpler, faster and end-to-end trainable. We release models, data, code and benchmarks under open licenses at https://hf.co/vidore .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper98
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen 等NeurIPS 2025 · 被引用 76 次
- Think Then Embed: Generative Context Improves Multimodal EmbeddingXuanming Cui, Jianpeng Cheng, Hong-You Chen, Satya Narayan Shukla 等ICLR 2026 · 被引用 41 次
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen 等ICLR 2026 · 被引用 40 次
- UME-R1: Exploring Reasoning-Driven Generative Multimodal EmbeddingsZhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou 等ICLR 2026 · 被引用 38 次
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb 等ACL 2025 · 被引用 33 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World ScenariosAntónio Loison, Quentin Macé, Antoine Edy, Victor Xing 等ACL 2026 · 被引用 15 次
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison 等ICML 2026 · 被引用 17 次
- ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained AlignmentHao Yang, Yifan Ji, Zhipeng Xu, Zhenghao Liu 等SIGIR 2026 · 被引用 4 次
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui 等ICLR 2025
- MMDocIR: Benchmarking Multimodal Retrieval for Long DocumentsKuicai Dong, Yujing Chang, Derrick-Goh-Xin Deik, Dexun Li 等EMNLP 2025 · 被引用 1 次
