Supporting Sensemaking of Large Language Model Outputs at Scale
Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, Elena L. Glassman
摘要
Large language models (LLMs) are capable of generating multiple responses to a single prompt, yet little effort has been expended to help end-users or system designers make use of this capability. In this paper, we explore how to present many LLM responses at once. We design five features, which include both pre-existing and novel methods for computing similarities and differences across textual documents, as well as how to render their outputs. We report on a controlled user study (n=24) and eight case studies evaluating these features and how they support users in different tasks. We find that the features support a wide variety of sensemaking tasks and even make tasks tractable that our participants previously considered to be too difficult to attempt. Finally, we present design guidelines to inform future explorations of new LLM interfaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran 等UIST 2024 · 被引用 143 次
- Textoshop: Interactions Inspired by Drawing Software to Facilitate Text EditingDamien Masson, Young-Ho Kim, Fanny ChevalierCHI 2025 · 被引用 38 次
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin 等CHI 2025 · 被引用 26 次
- AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose ToolsNathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown 等CHI 2025 · 被引用 26 次
它引用的顶会 Paper9
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 被引用 586 次
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language ModelsSangho Suh, Bryan Min, Srishti Palani, Haijun XiaUIST 2023 · 被引用 147 次
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg 等CHI 2024 · 被引用 141 次
- Graphologue: Exploring Large Language Model Responses with Interactive DiagramsPeiling Jiang, Jude Rayan, Steven P. Dow, Haijun XiaUIST 2023 · 被引用 135 次
- The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English WritersDaniel Buschek, Martin Zürn, Malin EibandCHI 2021 · 被引用 106 次
相关 Paper
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu 等IEEE VIS 2024 · 被引用 23 次
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li 等CHI 2024 · 被引用 143 次
- How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsHuichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn 等IEEE VIS 2024 · 被引用 12 次
- Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D ScenesJunlong Chen, Jens Grubert, Per Ola KristenssonIEEE VR 2025 · 被引用 9 次
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
