Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
Xingchen Zeng, Haichuan Lin, Yilin Ye, Wei Zeng
Abstract
Emerging multimodal large language models (MLLMs) exhibit great potential for chart question answering (CQA). Recent efforts primarily focus on scaling up training datasets (i.e., charts, data tables, and question-answer (QA) pairs) through data collection and synthesis. However, our empirical study on existing MLLMs and CQA datasets reveals notable gaps. First, current data collection and synthesis focus on data volume and lack consideration of fine-grained visual encodings and QA tasks, resulting in unbalanced data distribution divergent from practical CQA scenarios. Second, existing work follows the training recipe of the base MLLMs initially designed for natural images, under-exploring the adaptation to unique chart characteristics, such as rich text elements. To fill the gap, we propose a visualization-referenced instruction tuning approach to guide the training dataset enhancement and model development. Specifically, we propose a novel data engine to effectively filter diverse and high-quality data from existing datasets and subsequently refine and augment the data using LLM-based generation techniques to better align with practical QA tasks and visual encodings. Then, to facilitate the adaptation to chart characteristics, we utilize the enriched data to train an MLLM by unfreezing the vision encoder and incorporating a mixture-of-resolution adaptation strategy for enhanced fine-grained recognition. Experimental results validate the effectiveness of our approach. Even with fewer training examples, our model consistently outperforms state-of-the-art CQA models on established benchmarks. We also contribute a dataset split as a benchmark for future research. Source codes and datasets of this paper are available at https://github.com/zengxingchen/ChartQA-MLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8958780b-463b-4663-b3c5-39ba06d34e37Cited by top-tier papers19
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li et al.ICLR 2026 · 15 citations
- MultiVis-Agent: A Multi-Agent Framework with Logic Rules for Reliable and Comprehensive Cross-Modal Data VisualizationJinwei Lu, Yuanfeng Song, Chen Zhang, Raymond Chi-Wing WongSIGMOD 2026 · 14 citations
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 11 citations
- Protecting multimodal large language models against misleading visualizationsJonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna GurevychACL 2026 · 8 citations
- ChartCap: Mitigating Hallucination of Dense Chart CaptioningJunyoung Lim, Jaewoo Ahn, Gunhee KimICCV 2025 · 8 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- NovaChart: A Large-scale Dataset towards Chart Understanding and Generation of Multimodal Large Language ModelsLinmei Hu, Duokang Wang, Yiming Pan, Jifan Yu et al.ACM MM 2024 · 5 citations
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao et al.EMNLP 2025 · 6 citations
- Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackFatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, Gunhee KimEMNLP 2024 · 2 citations
- Effective Training Data Synthesis for Improving MLLM Chart UnderstandingYuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li et al.ICCV 2025 · 4 citations
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsXiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh et al.ICCV 2025 · 2 citations
