Effective Training Data Synthesis for Improving MLLM Chart Understanding
Yuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li, Gaowen Liu, Ali Payani, Yuan-Sen Ting, Liang Zheng
Abstract
Being able to effectively read scientific plots, or chart understanding, is a central part toward building effective agents for science. However, existing multimodal large language models (MLLMs), especially open-source ones, are still falling behind with a typical success rate of 30%-50% on challenging benchmarks. Previous studies on fine-tuning MLLMs with synthetic charts are often restricted by their inadequate similarity to the real charts, which could compromise model training and performance on complex real-world charts. In this study, we show that modularizing chart generation and diversifying visual details improves chart understanding capabilities. In particular, we design a five-step data synthesis pipeline, where we separate data and function creation for single plot generation, condition the generation of later subplots on earlier ones for multi-subplot figures, visually diversify the generated figures, filter out low quality data, and finally generate the question-answer (QA) pairs with GPT-4o. This approach allows us to streamline the generation of fine-tuning datasets and introduce the effective chart dataset (ECD), which contains 10k+ chart images and 300k+ QA pairs, covering 25 topics and featuring 250+ chart type combinations with high visual complexity. We show that ECD consistently improves the performance of various MLLMs on a range of real-world and synthetic test sets. Code, data and models are available at: https://github.com/yuweiyang-anu/ECD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 909e2e95-5f46-4937-a679-1fddcdcc9db3Cited by top-tier papers3
- Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-PlayQinsi Wang, Bo Liu, Tianyi Zhou, Jing Shi et al.ICLR 2026 · 24 citations
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal ReasoningShuoshuo Zhang, Yizhen Zhang, Jingjing Fu, Lei Song et al.CVPR 2026 · 3 citations
- R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?Jingyi Zhang, Tianyi Lin, Huanjin Yao, Xiang Lan et al.ICML 2026
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningAhmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque et al.EMNLP 2023 · 48 citations
- EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingMuye Huang, Han Lai, Xinyu Zhang, Wenjun Wu et al.AAAI 2025 · 30 citations
Related papers
- NovaChart: A Large-scale Dataset towards Chart Understanding and Generation of Multimodal Large Language ModelsLinmei Hu, Duokang Wang, Yiming Pan, Jifan Yu et al.ACM MM 2024 · 5 citations
- ChartGalaxy: A Dataset for Infographic Chart Understanding and GenerationZhen Li, Duan Li, Yukai Guo, Xinyuan Guo et al.ICLR 2026 · 16 citations
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez et al.CVPR 2026 · 7 citations
- Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction TuningXingchen Zeng, Haichuan Lin, Yilin Ye, Wei ZengIEEE VIS 2024 · 23 citations
- FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM ReasoningKaiwen Shi, Sichen Liu, Ziyue Lin, Hangrui Guo et al.ICLR 2026
