Latent Space Explanation by Intervention
Itai Gat, Guy Lorberbom, Idan Schwartz, Tamir Hazan
Abstract
The success of deep neural nets heavily relies on their ability to encode complex relations between their input and their output. While this property serves to fit the training data well, it also obscures the mechanism that drives prediction. This study aims to reveal hidden concepts by employing an intervention mechanism that shifts the predicted class based on discrete variational autoencoders. An explanatory model then visualizes the encoded information from any hidden layer and its corresponding intervened representation. By the assessment of differences between the original representation and the intervened representation, one can determine the concepts that can alter the class, hence providing interpretability. We demonstrate the effectiveness of our approach on CelebA, where we show various visualizations for bias in the data and suggest different interventions to reveal and change bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62dbdcad-2efa-4903-a92e-45c8c2cf0fe2Cited by top-tier papers4
- Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model AdaptationGuy Yariv, Itai Gat, Sagie Benaim, Lior Wolf et al.AAAI 2024 · 79 citations
- Layer Collaboration in the Forward-Forward AlgorithmGuy Lorberbom, Itai Gat, Yossi Adi, Alexander G. Schwing et al.AAAI 2024 · 22 citations
- Discriminative Class Tokens for Text-to-Image Diffusion ModelsIdan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer, Serge J. Belongie et al.ICCV 2023 · 13 citations
- A Functional Information Perspective on Model InterpretationItai Gat, Nitay Calderon, Roi Reichart, Tamir HazanICML 2022 · 6 citations
Builds on9
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
- A causal view of compositional zero-shot recognitionYuval Atzmon, Felix Kreuk, Uri Shalit, Gal ChechikNeurIPS 2020 · 163 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
Related papers
- CausalVAE: Disentangled Representation Learning via Neural Structural Causal ModelsMengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen et al.CVPR 2021
- Exploring the Latent Space of Autoencoders with Interventional AssaysFelix Leeb, Stefan Bauer, Michel Besserve, Bernhard SchölkopfNeurIPS 2022 · 26 citations
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 46 citations
- CF-OPT: Counterfactual Explanations for Structured PredictionGermain Vivier-Ardisson, Alexandre Forel, Axel Parmentier, Thibaut VidalICML 2024 · 3 citations
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng et al.CVPR 2021
