ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question Answering
Xiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang, Qin Zhang, Liang-Jie Zhang
摘要
Chart Question Answering (CQA) benchmarks are critical for evaluating Multimodal Large Language Models (MLLMs) on visual data reasoning. Existing benchmarks focus mainly on final-answer correctness, ignoring intermediate reasoning steps and the propagation of errors in multi-step processes. To address this, we introduce ChartR, a benchmark designed to assess both the accuracy and robustness of reasoning in chart-understanding tasks. Each question is decomposed into 4-10 sub-questions covering key reasoning types, and each chart includes four visually perturbed variants (blurred, noise-added, watermarkadded, annotation-removed) to systematically evaluate robustness. ChartR contains 200 base charts, 800 variants, 1,652 questions, and 8,260 image-question pairs. We further propose a comprehensive evaluation framework with eight metrics that evaluate reasoning-chain accuracy, robustness under visual perturbations, and enable analysis of potential error propagation patterns. Experiments on twelve MLLMs, including general-purpose and chartspecialized models, reveal low reasoning reliability, earlystep errors that may propagate, value extraction as the primary bottleneck, and sharp performance drops under perturbations, highlighting reliance on textual cues over true visual understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- VisText: A Benchmark for Semantically Rich Chart CaptioningBenny J. Tang, Angie W. Boggust, Arvind SatyanarayanACL 2023 · 被引用 44 次
- OpenCQA: Open-ended Question Answering with ChartsShankar Kantharaj, Xuan Long Do, Rixie Tiffany Ko Leong, Jia Qing Tan 等EMNLP 2022 · 被引用 26 次
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan 等EMNLP 2024 · 被引用 15 次
- OneChart: Purify the Chart Structural Extraction via One Auxiliary TokenJinyue Chen, Lingyu Kong, Haoran Wei, Chenglong Liu 等ACM MM 2024 · 被引用 13 次
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationCheng Yang, Chufan Shi, Yaxin Liu, Bo Shui 等ICLR 2025 · 被引用 3 次
相关 Paper
- DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsYujing Lu, Ling Zhong, Jing Yang, Weiming Li 等AAAI 2026
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 等EMNLP 2025 · 被引用 6 次
- DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense DocumentsZhuoran Yu, Le T Nguyen, Jaden Park, Xinyi Gu 等ICML 2026
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin 等EMNLP 2025 · 被引用 1 次
- Protecting multimodal large language models against misleading visualizationsJonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna GurevychACL 2026 · 被引用 8 次
