Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and Analysis
Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, Mennatallah El-Assady
摘要
A 23-year-old pregnant woman at 22 weeks gestation presents burning upon urination... Which of the following is the best treatment for this patient?
The answer is D. Nitrofurantoin CONTAIN "Nitrofurantoin" Outcome: A A 23-year-old pregnant woman at 22 weeks gestation presents burning upon urination... Which of the following is the best treatment for this patient? The answer is B. Ceftriaxone The Original Prompt and the LLM Response A Counterfactual Prompt and the LLM Response
Fig. 1: LLM Analyzer enables users to analyze and understand LLM behaviors through meaningful counterfactuals. From the user's input prototype sentences, such as a medical question in this example, the system generates meaningful segments for perturbation. (A) Users can interactively adjust their granularities and specify alternative segments for replacements, which are then used by the system to create meaningful counterfactuals. The system enables users to analyze the LLM by inspecting and interactively aggregating the counterfactual examples in a table-based visualization. (B) The table header shows the segments' text, dependencies, and feature attributions. (C) By grouping the counterfactuals by segments of interest, users can assess their joint influence on the predictions. (D) In addition to the statistical results, users could view concrete examples to understand the LLM's prediction in alternative scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 被引用 758 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg 等CHI 2024 · 被引用 141 次
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 被引用 118 次
相关 Paper
- Steering Semantic Data Processing With DocWranglerShreya Shankar, Bhavya Chopra, Mawil Hasan, Stephen Lee 等UIST 2025
- Medical Interpretability and Knowledge Maps of Large Language ModelsRazvan Marinescu, Victoria-Elisabeth Gruber, Diego Fajardo VargasICLR 2026 · 被引用 1 次
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 被引用 116 次
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?Daniel P. Jeong, Saurabh Garg, Zachary C. Lipton, Michael OberstEMNLP 2024 · 被引用 20 次
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual ExplanationsHarry Mayne, Ryan Othniel Kearns, Yushi Yang, Andrew M. Bean 等EMNLP 2025 · 被引用 9 次
