Probing the Visualization Literacy of Vision Language Models: The Good, the Bad, and the Ugly
Lianghan Dong, Anamaria Crisan
Abstract
Vision Language Models (VLMs) demonstrate promising chart comprehension capabilities. Yet, prior explorations of their visualization literacy have been limited to assessing their response correctness and fail to explore their internal reasoning. To address this gap, we adapted attention-guided class activation maps (AG-CAM) for VLMs, to visualize the influence and importance of input features (image and text) on model responses. Using this approach, we conducted an examination of four open-source (ChartGemma, Janus 1B and 7B, and LLaVA) and two closed-source (GPT-4o, Gemini) models comparing their performance and, for the open-source models, their AG-CAM results. Overall, we found that ChartGemma, a 3B parameter VLM fine-tuned for chart question-answering (QA), outperformed other open-source models and exhibited performance on par with significantly larger closed-source VLMs. We also found that VLMs exhibit spatial reasoning by accurately localizing key chart features, and semantic reasoning by associating visual elements with corresponding data values and query tokens. Our approach is the first to demonstrate the use of AG-CAM on early fusion VLM architectures, which are widely used, and for chart QA. We also show preliminary evidence that these results can align with human reasoning. Our promising open-source VLMs results pave the way for transparent and reproducible research in AI visualization literacy. Code and Supplemental Materials: https://osf.io/fp3rg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
Related papers
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 6 citations
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 40 citations
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan et al.EMNLP 2024 · 15 citations
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationCheng Yang, Chufan Shi, Yaxin Liu, Bo Shui et al.ICLR 2025 · 3 citations
- OneChart: Purify the Chart Structural Extraction via One Auxiliary TokenJinyue Chen, Lingyu Kong, Haoran Wei, Chenglong Liu et al.ACM MM 2024 · 13 citations
