DISSECT: Disentangled Simultaneous Explanations via Concept Traversals
Asma Ghandeharioun, Been Kim, Chun-Liang Li, Brendan Jou, Brian Eoff, Rosalind W. Picard
摘要
Explaining deep learning model inferences is a promising venue for scientific understanding, improving safety, uncovering hidden biases, evaluating fairness, and beyond, as argued by many scholars. One of the principal benefits of counterfactual explanations is allowing users to explore"what-if"scenarios through what does not and cannot exist in the data, a quality that many other forms of explanation such as heatmaps and influence functions are inherently incapable of doing. However, most previous work on generative explainability cannot disentangle important concepts effectively, produces unrealistic examples, or fails to retain relevant information. We propose a novel approach, DISSECT, that jointly trains a generator, a discriminator, and a concept disentangler to overcome such challenges using little supervision. DISSECT generates Concept Traversals (CTs), defined as a sequence of generated examples with increasing degrees of concepts that influence a classifier's decision. By training a generative model from a classifier's signal, DISSECT offers a way to discover a classifier's inherent"notion"of distinct concepts automatically rather than rely on user-predefined concepts. We show that DISSECT produces CTs that (1) disentangle several concepts, (2) are influential to a classifier's decision and are coupled to its reasoning due to joint training (3), are realistic, (4) preserve relevant information, and (5) are stable across similar inputs. We validate DISSECT on several challenging synthetic and realistic datasets where previous methods fall short of satisfying desirable criteria for interpretability and show that it performs consistently well and better than existing methods. Finally, we present experiments showing applications of DISSECT for detecting potential biases of a classifier and identifying spurious artifacts that impact predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 被引用 37 次
- Spuriosity Rankings: Sorting Data to Measure and Mitigate BiasesMazda Moayeri, Wenxiao Wang, Sahil Singla, Soheil FeiziNeurIPS 2023 · 被引用 19 次
- Concept Distillation: Leveraging Human-Centered Explanations for Model ImprovementAvani Gupta, Saurabh Saini, P. J. NarayananNeurIPS 2023 · 被引用 18 次
- Perceptual Pat: A Virtual Human Visual System for Iterative Visualization DesignSungbok Shin, Sanghyun Hong, Niklas ElmqvistCHI 2023 · 被引用 11 次
它引用的顶会 Paper9
- How Can I Explain This to You? An Empirical Study of Deep Neural Network Explanation MethodsJeya Vikranth Jeyakumar, Joseph Noor, Yu-Hsi Cheng, Luis Garcia 等NeurIPS 2020 · 被引用 173 次
- Weakly Supervised Disentanglement with GuaranteesRui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon 等ICLR 2020 · 被引用 148 次
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 被引用 116 次
- InfoGAN-CR and ModelCentrality: Self-supervised Model Training and Selection for Disentangling GANsZinan Lin, Kiran Koshy Thekumparampil, Giulia Fanti, Sewoong OhICML 2020 · 被引用 106 次
- Regularizing Black-box Models for Improved InterpretabilityGregory Plumb, Maruan Al-Shedivat, Ángel Alexander Cabrera, Adam Perer 等NeurIPS 2020 · 被引用 90 次
相关 Paper
- Beyond Trivial Counterfactual Explanations with Diverse Valuable ExplanationsPau Rodríguez, Massimo Caccia, Alexandre Lacoste, Lee Zamparo 等ICCV 2021 · 被引用 72 次
- Grounding Counterfactual Explanation of Image Classifiers to Textual Concept SpaceSiwon Kim, Jinoh Oh, Sungjin Lee, Seunghak Yu 等CVPR 2023
- CoLiDR: Concept Learning using Aggregated Disentangled RepresentationsSanchit Sinha, Guangzhi Xiong, Aidong ZhangKDD 2024
- Constructing Fair Latent Space for Intersection of Fairness and ExplainabilityHyungjun Joo, Hyeonggeun Han, Sehwan Kim, Sangwoo Hong 等AAAI 2025 · 被引用 2 次
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 被引用 11 次
