Write, Rank, or Rate: Comparing Methods for Studying Visualization Affordances
Chase Stokes, Kylie R. Lin, Cindy Xiong Bearfield
Abstract
A growing body of work on visualization affordances highlights how specific design choices shape reader takeaways from information visualizations. However, mapping the relationship between design choices and reader conclusions often requires labor-intensive crowdsourced studies, generating large corpora of free-response text for analysis. To address this challenge, we explored alternative scalable research methodologies to assess chart affordances. We test four elicitation methods from human-subject studies: free response, visualization ranking, conclusion ranking, and salience rating, and compare their effectiveness in eliciting reader interpretations of line charts, dot plots, and heatmaps. Overall, we find that while no method fully replicates affordances observed in free-response conclusions, combinations of ranking and rating methods can serve as an effective proxy at a broad scale. The two ranking methodologies were influenced by participant bias towards certain chart types and the comparison of suggested conclusions. Rating conclusion salience could not capture the specific variations between chart types observed in the other methods. To supplement this work, we present a case study with GPT-40, exploring the use of large language models (LLMs) to elicit human-like chart interpretations. This aligns with recent academic interest in leveraging LLMs as proxies for human participants to improve data collection and analysis efficiency. GPT-40 performed best as a human proxy for the salience rating methodology but suffered from severe constraints in other areas. Overall, the discrepancies in affordances we found between various elicitation methodologies, including GPT-40, highlight the importance of intentionally selecting and combining methods and evaluating trade-offs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c76aca95-e72b-4d22-b69f-6308a0d2515bBuilds on16
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 244 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che et al.ICSE 2023 · 107 citations
- Striking a Balance: Reader Takeaways and Preferences when Integrating Text and ChartsChase Stokes, Vidya Setlur, Bridget Cogley, Arvind Satyanarayan et al.IEEE VIS 2022 · 64 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
Related papers
- Eye of the Beholder: Towards Measuring Visualization ComplexityJohannes Ellemose, Niklas ElmqvistIEEE VIS 2025 · 2 citations
- How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsHuichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn et al.IEEE VIS 2024 · 12 citations
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 6 citations
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 40 citations
- DracoGPT: Extracting Visualization Design Preferences from Large Language ModelsHuichen Will Wang, Mitchell Gordon, Leilani Battle, Jeffrey HeerIEEE VIS 2024 · 19 citations
