How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying Layouts
Huichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn, Cindy Xiong Bearfield
Abstract
Large Language Models (LLMs) have been adopted for a variety of visualizations tasks, but how far are we from perceptually aware LLMs that can predict human takeaways? Graphical perception literature has shown that human chart takeaways are sensitive to visualization design choices, such as spatial layouts. In this work, we examine the extent to which LLMs exhibit such sensitivity when generating takeaways, using bar charts with varying spatial layouts as a case study. We conducted three experiments and tested four common bar chart layouts: vertically juxtaposed, horizontally juxtaposed, overlaid, and stacked. In Experiment 1, we identified the optimal configurations to generate meaningful chart takeaways by testing four LLMs, two temperature settings, nine chart specifications, and two prompting strategies. We found that even state-of-the-art LLMs struggled to generate semantically diverse and factually accurate takeaways. In Experiment 2, we used the optimal configurations to generate 30 chart takeaways each for eight visualizations across four layouts and two datasets in both zero-shot and one-shot settings. Compared to human takeaways, we found that the takeaways LLMs generated often did not match the types of comparisons made by humans. In Experiment 3, we examined the effect of chart context and data on LLM takeaways. We found that LLMs, unlike humans, exhibited variation in takeaway comparison types for different bar charts using the same bar layout. Overall, our case study evaluates the ability of LLMs to emulate human interpretations of data and points to challenges and opportunities in using LLMs to predict human chart takeaways.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a809926-d8b6-422b-84de-139a49bd0a77Cited by top-tier papers4
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 11 citations
- Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging SystemShuyu Shen, Sirong Lu, Leixian Shen, Yuyu LuoCHI 2026 · 2 citations
- Write, Rank, or Rate: Comparing Methods for Studying Visualization AffordancesChase Stokes, Kylie R. Lin, Cindy Xiong BearfieldIEEE VIS 2025 · 2 citations
- A Rigorous Behavior Assessment of CNNs Using a Data-Domain Sampling RegimeShuning Jiang, Wei-Lun Chao, Daniel Haehn, Hanspeter Pfister et al.IEEE VIS 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet et al.NeurIPS 2024 · 271 citations
- Accessible Visualization via Natural Language Descriptions: A Four-Level Model of Semantic ContentAlan Lundgard, Arvind SatyanarayanIEEE VIS 2021 · 141 citations
Related papers
- Visual Arrangements of Bar Charts Influence Comparisons in Viewer TakeawaysCindy Xiong, Vidya Setlur, Benjamin Bach, Eunyee Koh et al.IEEE VIS 2021 · 40 citations
- How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?Leo Yu-Ho Lo, Huamin QuIEEE VIS 2024 · 24 citations
- Doc2Chart: Intent-Driven Zero-Shot Chart Generation from DocumentsAkriti Jain, Pritika Ramu, Aparna Garimella, Apoorv SaxenaEMNLP 2025
- Protecting multimodal large language models against misleading visualizationsJonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna GurevychACL 2026 · 8 citations
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin et al.EMNLP 2025 · 1 citation
