Supporting Sensemaking of Large Language Model Outputs at Scale
Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, Elena L. Glassman
Abstract
Large language models (LLMs) are capable of generating multiple responses to a single prompt, yet little effort has been expended to help end-users or system designers make use of this capability. In this paper, we explore how to present many LLM responses at once. We design five features, which include both pre-existing and novel methods for computing similarities and differences across textual documents, as well as how to render their outputs. We report on a controlled user study (n=24) and eight case studies evaluating these features and how they support users in different tasks. We find that the features support a wide variety of sensemaking tasks and even make tasks tractable that our participants previously considered to be too difficult to attempt. Finally, we present design guidelines to inform future explorations of new LLM interfaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c64d54d7-3bbb-4130-8af5-ebb8d8085021Cited by top-tier papers26
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- Textoshop: Interactions Inspired by Drawing Software to Facilitate Text EditingDamien Masson, Young-Ho Kim, Fanny ChevalierCHI 2025 · 38 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
- AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose ToolsNathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown et al.CHI 2025 · 26 citations
Builds on9
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 586 citations
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language ModelsSangho Suh, Bryan Min, Srishti Palani, Haijun XiaUIST 2023 · 147 citations
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg et al.CHI 2024 · 141 citations
- Graphologue: Exploring Large Language Model Responses with Interactive DiagramsPeiling Jiang, Jude Rayan, Steven P. Dow, Haijun XiaUIST 2023 · 135 citations
- The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English WritersDaniel Buschek, Martin Zürn, Malin EibandCHI 2021 · 106 citations
Related papers
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li et al.CHI 2024 · 143 citations
- How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsHuichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn et al.IEEE VIS 2024 · 12 citations
- Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D ScenesJunlong Chen, Jens Grubert, Per Ola KristenssonIEEE VR 2025 · 9 citations
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
