ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic Rules
Zhi-Qi Cheng, Qi Dai, Alexander G. Hauptmann
摘要
Charts are a powerful tool for visually conveying complex data, but their comprehension poses a challenge due to the diverse chart types and intricate components. Existing chart comprehension methods suffer from either heuristic rules or an over-reliance on OCR systems, resulting in suboptimal performance. To address these issues, we present ChartReader, a unified framework that seamlessly integrates chart derendering and comprehension tasks. Our approach includes a transformer-based chart component detection module and an extended pre-trained vision-language model for chart-to-X tasks. By learning the rules of charts automatically from annotated datasets, our approach eliminates the need for manual rule-making, reducing effort and enhancing accuracy. We also introduce a data variable replacement technique and extend the input and position embeddings of the pre-trained model for cross-task training. We evaluate ChartReader on Chart-to-Table , ChartQA, and Chart-to-Text tasks, demonstrating its superiority over existing methods. Our proposed framework can significantly reduce the manual effort involved in chart analysis, providing a step towards a universal chart understanding model. Moreover, our approach offers opportunities for plug-and-play integration with mainstream LLMs such as T5 and TaPas, extending their capability to chart comprehension tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingMuye Huang, Han Lai, Xinyu Zhang, Wenjun Wu 等AAAI 2025 · 被引用 30 次
- Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction TuningXingchen Zeng, Haichuan Lin, Yilin Ye, Wei ZengIEEE VIS 2024 · 被引用 23 次
- VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution ReasoningMuye Huang, Lingling Zhang, Han Lai, Wenjun Wu 等AAAI 2025 · 被引用 7 次
- Union Is Strength! Unite the Power of LLMs and MLLMs for Chart Question AnsweringJiapeng Liu, Liang Li, Shihao Rao, Xiyan Gao 等AAAI 2025 · 被引用 3 次
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper10
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju 等CVPR 2022 · 被引用 82 次
- STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question AnsweringHrituraj Singh, Sumit ShekharEMNLP 2020 · 被引用 42 次
相关 Paper
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningAhmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque 等EMNLP 2023 · 被引用 48 次
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 等EMNLP 2025 · 被引用 6 次
- ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question AnsweringRachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra GaneshACL 2026 · 被引用 5 次
- MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart DerenderingFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang 等ACL 2023 · 被引用 38 次
- Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart UnderstandingJianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu 等ACL 2026
