Write, Rank, or Rate: Comparing Methods for Studying Visualization Affordances
Chase Stokes, Kylie R. Lin, Cindy Xiong Bearfield
摘要
A growing body of work on visualization affordances highlights how specific design choices shape reader takeaways from information visualizations. However, mapping the relationship between design choices and reader conclusions often requires labor-intensive crowdsourced studies, generating large corpora of free-response text for analysis. To address this challenge, we explored alternative scalable research methodologies to assess chart affordances. We test four elicitation methods from human-subject studies: free response, visualization ranking, conclusion ranking, and salience rating, and compare their effectiveness in eliciting reader interpretations of line charts, dot plots, and heatmaps. Overall, we find that while no method fully replicates affordances observed in free-response conclusions, combinations of ranking and rating methods can serve as an effective proxy at a broad scale. The two ranking methodologies were influenced by participant bias towards certain chart types and the comparison of suggested conclusions. Rating conclusion salience could not capture the specific variations between chart types observed in the other methods. To supplement this work, we present a case study with GPT-40, exploring the use of large language models (LLMs) to elicit human-like chart interpretations. This aligns with recent academic interest in leveraging LLMs as proxies for human participants to improve data collection and analysis efficiency. GPT-40 performed best as a human proxy for the salience rating methodology but suffered from severe constraints in other areas. Overall, the discrepancies in affordances we found between various elicitation methodologies, including GPT-40, highlight the importance of intentionally selecting and combining methods and evaluating trade-offs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 被引用 244 次
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover 等IEEE VIS 2022 · 被引用 191 次
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
- Striking a Balance: Reader Takeaways and Preferences when Integrating Text and ChartsChase Stokes, Vidya Setlur, Bridget Cogley, Arvind Satyanarayan 等IEEE VIS 2022 · 被引用 64 次
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren 等IEEE VIS 2024 · 被引用 44 次
相关 Paper
- Eye of the Beholder: Towards Measuring Visualization ComplexityJohannes Ellemose, Niklas ElmqvistIEEE VIS 2025 · 被引用 2 次
- How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsHuichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn 等IEEE VIS 2024 · 被引用 12 次
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 被引用 6 次
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 被引用 40 次
- DracoGPT: Extracting Visualization Design Preferences from Large Language ModelsHuichen Will Wang, Mitchell Gordon, Leilani Battle, Jeffrey HeerIEEE VIS 2024 · 被引用 19 次
