Promises and Pitfalls: Using Large Language Models to Generate Visualization Items
Yuan Cui, Lily W. Ge, Yiren Ding, Lane Harrison, Fumeng Yang, Matthew Kay
摘要
Visualization items-factual questions about visualizations that ask viewers to accomplish visualization tasks-are regularly used in the field of information visualization as educational and evaluative materials. For example, researchers of visualization literacy require large, diverse banks of items to conduct studies where the same skill is measured repeatedly on the same participants. Yet, generating a large number of high-quality, diverse items requires significant time and expertise. To address the critical need for a large number of diverse visualization items in education and research, this paper investigates the potential for large language models (LLMS) to automate the generation of multiple-choice visualization items. Through an iterative design process, we develop the VILA (Visualization Items Generated by Large LAnguage Models) pipeline, for efficiently generating visualization items that measure people's ability to accomplish visualization tasks. We use the VILA pipeline to generate 1,404 candidate items across 12 chart types and 13 visualization tasks. In collaboration with 11 visualization experts, we develop an evaluation rulebook which we then use to rate the quality of all candidate items. The result is the VILA bank of 1, 100 items. From this evaluation, we also identify and classify current limitations of the VILA pipeline, and discuss the role of human oversight in ensuring quality. In addition, we demonstrate an application of our work by creating a visualization literacy test, VILA-VLAT, which measures people's ability to complete a diverse set of tasks on various types of visualizations; comparing it to the existing VLAT, VILA-VLAT shows moderate to high convergent validity (R = 0.70). Lastly, we discuss the application areas of the VILA pipeline and the VILA bank and provide practical recommendations for their use. All supplemental materials are available at https://osf.io/ysrhq/.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- ReVISit 2: A Full Experiment Life Cycle User Study FrameworkZach Cutler, Jack Wilburn, Hilson Shrestha, Yiren Ding 等IEEE VIS 2025 · 被引用 25 次
- AVEC: An Assessment of Visual Encoding Ability in Visualization ConstructionLily W. Ge, Yuan Cui, Matthew KayCHI 2025 · 被引用 8 次
- EncQA: Benchmarking Vision-Language Models on Visual Encodings for ChartsKushin Mukherjee, Donghao Ren, Dominik Moritz, Yannick AssogbaIEEE VIS 2025 · 被引用 3 次
- Embodied Natural Language Interaction (NLI): Speech Input Patterns in Immersive AnalyticsHyemi Song, Matthew Johnson, Kirsten Whitley, Eric Krokos 等IEEE VIS 2025 · 被引用 2 次
- An Autoethnography on Visualization Literacy: A Wicked Measurement ProblemLily W. Ge, Anne-Flore Cabouat, Karen Bonilla, Yuan Cui 等IEEE VIS 2025 · 被引用 2 次
相关 Paper
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 被引用 40 次
- CALVI: Critical Thinking Assessment for Literacy in VisualizationsLily W. Ge, Yuan Cui, Matthew KayCHI 2023 · 被引用 68 次
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 被引用 6 次
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 被引用 1 次
- Adaptive Assessment of Visualization LiteracyYuan Cui, Lily W. Ge, Yiren Ding, Fumeng Yang 等IEEE VIS 2023 · 被引用 26 次
