An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy Tasks
Alexander Bendeck, John T. Stasko
摘要
Large Language Models (LLMs) like GPT-4 which support multimodal input (i.e., prompts containing images in addition to text) have immense potential to advance visualization research. However, many questions exist about the visual capabilities of such models, including how well they can read and interpret visually represented data. In our work, we address this question by evaluating the GPT-4 multimodal LLM using a suite of task sets meant to assess the model's visualization literacy. The task sets are based on existing work in the visualization community addressing both automated chart question answering and human visualization literacy across multiple settings. Our assessment finds that GPT-4 can perform tasks such as recognizing trends and extreme values, and also demonstrates some understanding of visualization design best-practices. By contrast, GPT-4 struggles with simple value retrieval when not provided with the original dataset, lacks the ability to reliably distinguish between colors in charts, and occasionally suffers from hallucination and inconsistency. We conclude by reflecting on the model's strengths and weaknesses as well as the potential utility of models like GPT-4 for future visualization research. We also release all code, stimuli, and results for the task sets at the following link: https://doi.org/10.17605/OSF.IO/F39J6.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper12
- Protecting multimodal large language models against misleading visualizationsJonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna GurevychACL 2026 · 被引用 8 次
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 被引用 6 次
- MisVisFix: An Interactive Dashboard for Detecting, Explaining, and Correcting Misleading Visualizations using Large Language ModelsAmit Kumar Das, Klaus MuellerIEEE VIS 2025 · 被引用 5 次
- Is this chart lying to me? Automating the detection of misleading visualizationsJonathan Tonglet, Jan Zimny, Tinne Tuytelaars, Iryna GurevychACL 2026 · 被引用 4 次
- EncQA: Benchmarking Vision-Language Models on Visual Encodings for ChartsKushin Mukherjee, Donghao Ren, Dominik Moritz, Yannick AssogbaIEEE VIS 2025 · 被引用 3 次
相关 Paper
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 被引用 1 次
- How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?Leo Yu-Ho Lo, Huamin QuIEEE VIS 2024 · 被引用 24 次
- Probing the Visualization Literacy of Vision Language Models: The Good, the Bad, and the UglyLianghan Dong, Anamaria CrisanIEEE VIS 2025 · 被引用 1 次
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 被引用 1 次
- Promises and Pitfalls: Using Large Language Models to Generate Visualization ItemsYuan Cui, Lily W. Ge, Yiren Ding, Lane Harrison 等IEEE VIS 2024 · 被引用 17 次
