ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
António Loison, Quentin Macé, Antoine Edy, Victor Xing, Tom Balough, Gabriel de Souza Pereira Moreira, Bo Liu, Manuel Faysse, Céline Hudelot, Gautier Viaud
Abstract
Retrieval-Augmented Generation (RAG) pipelines must address challenges beyond simple single-document retrieval, such as interpreting visual elements (tables, charts, images), synthesizing information across documents, and providing accurate source grounding. Existing benchmarks fail to capture this complexity, often focusing on textual data, single-document comprehension, or evaluating retrieval and generation in isolation. We introduce ViDoRe V3, a comprehensive multimodal RAG benchmark featuring multi-type queries over visually rich document corpora. It covers 10 datasets across diverse professional domains, comprising 26,000 document pages paired with 3,099 human-verified queries, each available in 6 languages. Through 12,000 hours of human annotation effort, we provide high-quality annotations for retrieval relevance, bounding box localization, and verified reference answers. Our evaluation of state-of-the-art RAG pipelines reveals that visual retrievers outperform textual ones, late-interaction models and textual reranking substantially improve performance, and hybrid or purely visual contexts enhance answer generation quality. However, current models still struggle with non-textual elements, open-ended queries, and fine-grained visual grounding. To encourage progress in addressing these challenges, the benchmark is released under a commercially permissive license 1 . * Equal contribution † Work done while at Illuin Technology ‡ Contact emails 1 https://hf.co/vidore Query
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a2cd1b8-1f5a-4441-9a55-623aed4b5f22Cited by top-tier papers2
- LEMUR: Learned Multi-Vector RetrievalElias Jääsaari, Ville Hyvönen, Teemu RoosICML 2026 · 3 citations
- Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document CollectionsLukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha et al.ICML 2026
Builds on7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 152 citations
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb et al.ACL 2025 · 33 citations
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison et al.ICML 2026 · 17 citations
- ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning AgentsQiuchen Wang, Ruixue Ding, Zehui Chen, Weiqi Wu et al.EMNLP 2025 · 6 citations
Related papers
- ColPali: Efficient Document Retrieval with Vision Language ModelsManuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani et al.ICLR 2025 · 4 citations
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial DomainSuifeng Zhao, Zhuoran Jin, Sujian Li, Jun GaoEMNLP 2025 · 1 citation
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen et al.AAAI 2026
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan et al.ACL 2026 · 7 citations
- M3Retrieve: Benchmarking Multimodal Retrieval for MedicineArkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa et al.EMNLP 2025
