Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents
Yuming Yang, Jiang Zhong, Li Jin, Xiao Sun, Jingwang Huang, Jinpeng Gao, Qing Liu, Yang Bai, Jingyuan Zhang, Rui Jiang, Qin Lei, Kaiwen Wei
Abstract
Multimodal Retrieval-Augmented Generation (MRAG) enhances reasoning capabilities by integrating external knowledge. However, existing benchmarks primarily focus on simple image-text interactions, overlooking complex visual formats like charts that are prevalent in real-world applications. In this work, we introduce a novel task, Chart-based MRAG, to address this limitation. To generate high-quality evaluation samples, we propose CHARGE (CHARt-based document question-answering GEneration), a semi-automatic framework for generating evaluation samples through multimodal keypoint extraction, knowledge graph construction, and qa pair synthesis. By combining CHARGE with expert validation, we construct Chart-MRAG Bench, a comprehensive benchmark for chart-based MRAG evaluation, featuring 4,738 question-answering pairs across 8 domains from real-world documents. Our experiments reveal three critical limitations in current approaches: (1) unified multimodal embedding retrieval methods struggles in chartbased scenarios, (2) even with ground-truth retrieval, state-of-the-art Multimodal Large Language Models (MLLMs) achieve only 71.15% Correctness and 80.74% Coverage scores, and (3) Widely-used MLLMs demonstrate consistent text-over-visual modality bias. These findings highlight great challenges in processing information-dense visual formats. The dataset and code are available at Chart-MRAG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47062ff2-85a8-4b78-8dc8-068ea4eb8889Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga et al.EMNLP 2022 · 89 citations
- Unifying Multimodal Retrieval via Document Screenshot EmbeddingXueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen et al.EMNLP 2024 · 13 citations
Related papers
- MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal ModelsWenbo Hu, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz et al.ICLR 2025 · 1 citation
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang et al.CVPR 2026
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao et al.EMNLP 2025 · 6 citations
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationKaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang et al.WWW 2026
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen et al.AAAI 2026
