Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and Analysis
Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, Mennatallah El-Assady
Abstract
A 23-year-old pregnant woman at 22 weeks gestation presents burning upon urination... Which of the following is the best treatment for this patient?
The answer is D. Nitrofurantoin CONTAIN "Nitrofurantoin" Outcome: A A 23-year-old pregnant woman at 22 weeks gestation presents burning upon urination... Which of the following is the best treatment for this patient? The answer is B. Ceftriaxone The Original Prompt and the LLM Response A Counterfactual Prompt and the LLM Response
Fig. 1: LLM Analyzer enables users to analyze and understand LLM behaviors through meaningful counterfactuals. From the user's input prototype sentences, such as a medical question in this example, the system generates meaningful segments for perturbation. (A) Users can interactively adjust their granularities and specify alternative segments for replacements, which are then used by the system to create meaningful counterfactuals. The system enables users to analyze the LLM by inspecting and interactively aggregating the counterfactual examples in a table-based visualization. (B) The table header shows the segments' text, dependencies, and feature attributions. (C) By grouping the counterfactuals by segments of interest, users can assess their joint influence on the predictions. (D) In addition to the statistical results, users could view concrete examples to understand the LLM's prediction in alternative scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg et al.CHI 2024 · 141 citations
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 118 citations
Related papers
- Steering Semantic Data Processing With DocWranglerShreya Shankar, Bhavya Chopra, Mawil Hasan, Stephen Lee et al.UIST 2025
- Medical Interpretability and Knowledge Maps of Large Language ModelsRazvan Marinescu, Victoria-Elisabeth Gruber, Diego Fajardo VargasICLR 2026 · 1 citation
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 116 citations
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?Daniel P. Jeong, Saurabh Garg, Zachary C. Lipton, Michael OberstEMNLP 2024 · 20 citations
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual ExplanationsHarry Mayne, Ryan Othniel Kearns, Yushi Yang, Andrew M. Bean et al.EMNLP 2025 · 9 citations
