: Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language Models
Xingbo Wang, Renfei Huang, Zhihua Jin, Tianqing Fang, Huamin Qu
Abstract
Visualization system Commonsense QA Model editing Model understanding Model prediction Behavior diagnosis Commonsense knowledge extraction Multilevel visualization where do adults use glue sticks? A. classroom B. office C. desk drawer adult glue stick office work use atLocation commonsense Do NLP models have commonsense? Users NLP model
Fig. 1: Since commonsense knowledge is not explicitly stated, it is challenging to conduct a scalable analysis of what commonsense knowledge NLP models do (not) learn. We employ a knowledge graph to derive implicit commonsense in the model input as context information. Then, we use it to align model behavior with human reasoning through multi-level interactive visualizations. Thereafter, users can understand, diagnose, and edit specific knowledge areas where models do not perform well.
Abstract-Recently, large pretrained language models have achieved compelling performance on commonsense benchmarks. Nevertheless, it is unclear what commonsense knowledge the models learn and whether they solely exploit spurious patterns. Feature attributions are popular explainability techniques that identify important input concepts for model outputs. However, commonsense knowledge tends to be implicit and rarely explicitly presented in inputs. These methods cannot infer models' implicit reasoning over mentioned concepts. We present CommonsenseVIS, a visual explanatory system that utilizes external commonsense knowledge bases to contextualize model behavior for commonsense question-answering. Specifically, we extract relevant commonsense knowledge in inputs as references to align model behavior with human knowledge. Our system features multi-level visualization and interactive model probing and editing for different concepts and their underlying relations. Through a user study, we show that CommonsenseVIS helps NLP experts conduct a systematic and scalable visual analysis of models' relational reasoning over concepts in different situations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7c118cd-b75c-4e79-8300-655796474bffCited by top-tier papers5
- Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in EducationLin Gao, Jing Lu, Zekai Shao, Ziyue Lin et al.IEEE VIS 2024 · 27 citations
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented GenerationSizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin MaUIST 2025 · 6 citations
- Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and AnalysisFurui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst et al.IEEE VIS 2025 · 5 citations
- Crowdsourced Think-Aloud StudiesZach Cutler, Lane Harrison, Carolina Nobre, Alexander LexCHI 2025 · 4 citations
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li et al.IEEE VIS 2025 · 3 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
Related papers
- Toward Explainable Physical Audiovisual Commonsense ReasoningDaoming Zong, Chaoyue Ding, Kaitao ChenACM MM 2024
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck et al.ACL 2022
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume et al.EMNLP 2022 · 34 citations
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma et al.NeurIPS 2024 · 23 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
