Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
Xingchen Zeng, Haichuan Lin, Yilin Ye, Wei Zeng
摘要
Emerging multimodal large language models (MLLMs) exhibit great potential for chart question answering (CQA). Recent efforts primarily focus on scaling up training datasets (i.e., charts, data tables, and question-answer (QA) pairs) through data collection and synthesis. However, our empirical study on existing MLLMs and CQA datasets reveals notable gaps. First, current data collection and synthesis focus on data volume and lack consideration of fine-grained visual encodings and QA tasks, resulting in unbalanced data distribution divergent from practical CQA scenarios. Second, existing work follows the training recipe of the base MLLMs initially designed for natural images, under-exploring the adaptation to unique chart characteristics, such as rich text elements. To fill the gap, we propose a visualization-referenced instruction tuning approach to guide the training dataset enhancement and model development. Specifically, we propose a novel data engine to effectively filter diverse and high-quality data from existing datasets and subsequently refine and augment the data using LLM-based generation techniques to better align with practical QA tasks and visual encodings. Then, to facilitate the adaptation to chart characteristics, we utilize the enriched data to train an MLLM by unfreezing the vision encoder and incorporating a mixture-of-resolution adaptation strategy for enhanced fine-grained recognition. Experimental results validate the effectiveness of our approach. Even with fewer training examples, our model consistently outperforms state-of-the-art CQA models on established benchmarks. We also contribute a dataset split as a benchmark for future research. Source codes and datasets of this paper are available at https://github.com/zengxingchen/ChartQA-MLLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li 等ICLR 2026 · 被引用 15 次
- MultiVis-Agent: A Multi-Agent Framework with Logic Rules for Reliable and Comprehensive Cross-Modal Data VisualizationJinwei Lu, Yuanfeng Song, Chen Zhang, Raymond Chi-Wing WongSIGMOD 2026 · 被引用 14 次
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 被引用 11 次
- Protecting multimodal large language models against misleading visualizationsJonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna GurevychACL 2026 · 被引用 8 次
- ChartCap: Mitigating Hallucination of Dense Chart CaptioningJunyoung Lim, Jaewoo Ahn, Gunhee KimICCV 2025 · 被引用 8 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
相关 Paper
- NovaChart: A Large-scale Dataset towards Chart Understanding and Generation of Multimodal Large Language ModelsLinmei Hu, Duokang Wang, Yiming Pan, Jifan Yu 等ACM MM 2024 · 被引用 5 次
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 等EMNLP 2025 · 被引用 6 次
- Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackFatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, Gunhee KimEMNLP 2024 · 被引用 2 次
- Effective Training Data Synthesis for Improving MLLM Chart UnderstandingYuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li 等ICCV 2025 · 被引用 4 次
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsXiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh 等ICCV 2025 · 被引用 2 次
