NovaChart: A Large-scale Dataset towards Chart Understanding and Generation of Multimodal Large Language Models
Linmei Hu, Duokang Wang, Yiming Pan, Jifan Yu, Yingxia Shao, Chong Feng, Liqiang Nie
摘要
Multimodal Large Language Models (MLLMs) have shown significant potential for chart understanding and generation. However, they are still far from achieving the desired effectiveness in practical applications. This could be due to the limitations of the used training chart data. Existing chart datasets suffer from scarcity of chart types, limited coverage of tasks, and insufficient scalability, making them incapable of effectively enhancing the chart-related capabilities of MLLMs. To tackle these obstacles, we construct NovaChart, a large-scale dataset for chart understanding and generation of MLLMs. NovaChart contains 47K high-resolution chart images and 856K chart-related instructions, covering 18 different chart types and 15 unique tasks of chart understanding and generation. To build NovaChart, we propose a data generation engine for metadata curation, chart visualization and instruction formulation. Chart metadata in NovaChart contains detailed annotations, i.e., data points, visual elements, source data and the visualization code of every chart. This additional information endows NovaChart with considerable scalability, as it can facilitate the extension of chart instruction data to a larger scale and greater diversity. We utilize NovaChart to train several open-source MLLMs. Experimental results demonstrate NovaChart empowers MLLMs with stronger capabilities in 15 chart understanding and generation tasks by a large-margin (35.47%-619.47%), bringing them a step closer to smart chart assistants. Our dataset is now available at https://github.com/Elucidator-V/NovaChart.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- ChartGalaxy: A Dataset for Infographic Chart Understanding and GenerationZhen Li, Duan Li, Yukai Guo, Xinyuan Guo 等ICLR 2026 · 被引用 16 次
- Effective Training Data Synthesis for Improving MLLM Chart UnderstandingYuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li 等ICCV 2025 · 被引用 4 次
- Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense ChartsHongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu 等CVPR 2026 · 被引用 1 次
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source SuitesZhenxin Lei, Zhangwei Gao, Changyao Tian, Erfei Cui 等ICLR 2026 · 被引用 1 次
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning BenchmarkYunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li 等ICML 2025
相关 Paper
- Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction TuningXingchen Zeng, Haichuan Lin, Yilin Ye, Wei ZengIEEE VIS 2024 · 被引用 23 次
- ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code GenerationXuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen 等ACL 2025 · 被引用 62 次
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez 等CVPR 2026 · 被引用 7 次
- From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang 等ACL 2026 · 被引用 5 次
- Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackFatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, Gunhee KimEMNLP 2024 · 被引用 2 次
