Lune

CVPR2026Top-tier venue

ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question Answering

Xiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang, Qin Zhang, Liang-Jie Zhang

2026Year

Abstract

Chart Question Answering (CQA) benchmarks are critical for evaluating Multimodal Large Language Models (MLLMs) on visual data reasoning. Existing benchmarks focus mainly on final-answer correctness, ignoring intermediate reasoning steps and the propagation of errors in multi-step processes. To address this, we introduce ChartR, a benchmark designed to assess both the accuracy and robustness of reasoning in chart-understanding tasks. Each question is decomposed into 4-10 sub-questions covering key reasoning types, and each chart includes four visually perturbed variants (blurred, noise-added, watermarkadded, annotation-removed) to systematically evaluate robustness. ChartR contains 200 base charts, 800 variants, 1,652 questions, and 8,260 image-question pairs. We further propose a comprehensive evaluation framework with eight metrics that evaluate reasoning-chain accuracy, robustness under visual perturbations, and enable analysis of potential error propagation patterns. Experiments on twelve MLLMs, including general-purpose and chartspecialized models, reveal low reasoning reliability, earlystep errors that may propagate, value extraction as the primary bottleneck, and sharp performance drops under perturbations, highlighting reliance on textual cues over true visual understanding.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e419f22d-432d-424d-96ec-68869a1518ac

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines