UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Ahmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque, Shafiq Joty
摘要
Charts are widely used for data analysis, providing visual representations and insights into complex data. To facilitate chart-based data analysis using natural language, several downstream tasks have been introduced recently such as chart question answering and chart summarization. However, existing methods for these tasks often rely on pretraining on language or vision-language tasks, neglecting the explicit modeling of chart structures (e.g., how chart elements are related to each other). To address this, we first build a large corpus of charts covering diverse topics and visual styles. We then present UniChart, a pretrained model for chart comprehension and reasoning. UniChart encodes the relevant text, data, and visual elements of charts and then uses a chart-grounded text decoder for text generation. We propose several chart-specific pretraining tasks that include: (i) low-level tasks to extract the visual elements (e.g., bars, lines) and data from charts, and (ii) high-level tasks to acquire chart understanding and reasoning skills. Our experiments demonstrate that pretraining UniChart on a large corpus with chart-specific objectives, followed by fine-tuning, yields state-of-the-art performance on four downstream tasks. Moreover, our model exhibits superior generalizability to unseen chart corpus, surpassing previous approaches that lack chart-specific objectives and utilize limited chart resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code GenerationXuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen 等ACL 2025 · 被引用 62 次
- EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingMuye Huang, Han Lai, Xinyu Zhang, Wenjun Wu 等AAAI 2025 · 被引用 30 次
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsYi Zhang, Bolin Ni, Xin-Sheng Chen, Hengrui Zhang 等ICLR 2026 · 被引用 30 次
- How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?Leo Yu-Ho Lo, Huamin QuIEEE VIS 2024 · 被引用 24 次
它引用的顶会 Paper32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic RulesZhi-Qi Cheng, Qi Dai, Alexander G. HauptmannICCV 2023 · 被引用 32 次
- Doc2Chart: Intent-Driven Zero-Shot Chart Generation from DocumentsAkriti Jain, Pritika Ramu, Aparna Garimella, Apoorv SaxenaEMNLP 2025
- STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question AnsweringHrituraj Singh, Sumit ShekharEMNLP 2020 · 被引用 42 次
- MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart DerenderingFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang 等ACL 2023 · 被引用 38 次
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez 等CVPR 2026 · 被引用 7 次
