A Disentangling Invertible Interpretation Network for Explaining Latent Representations
Patrick Esser, Robin Rombach, Björn Ommer
Abstract
Neural networks have greatly boosted performance in computer vision by learning powerful representations of input data. The drawback of end-to-end training for maximal overall performance are black-box models whose hidden representations are lacking interpretability: Since distributed coding is optimal for latent layers to improve their robustness, attributing meaning to parts of a hidden feature vector or to individual neurons is hindered. We formulate interpretation as a translation of hidden representations onto semantic concepts that are comprehensible to the user. The mapping between both domains has to be bijective so that semantic modifications in the target domain correctly alter the original representation. The proposed invertible interpretation network can be transparently applied on top of existing architectures with no need to modify or retrain them. Consequently, we translate an original representation to an equivalent yet interpretable one and backwards without affecting the expressiveness and performance of the original. The invertible interpretation network disentangles the hidden representation into separate, semantically meaningful concepts. Moreover, we present an efficient approach to define semantic concepts by only sketching two images and also an unsupervised strategy. Experimental evaluation demonstrates the wide applicability to interpretation of existing classification and image generation networks as well as to semantically guided image manipulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef82b78e-53ac-47bf-86e4-a88389f22bc9Cited by top-tier papers22
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image SynthesisPatrick Esser, Robin Rombach, Andreas Blattmann, Björn OmmerNeurIPS 2021 · 187 citations
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald et al.ICCV 2021 · 181 citations
- Deep Digging into the Generalization of Self-Supervised Monocular Depth EstimationJinwoo Bae, Sungho Moon, Sunghoon ImAAAI 2023 · 127 citations
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- Shape or Texture: Understanding Discriminative Features in CNNsMd. Amirul Islam, Matthew Kowal, Patrick Esser, Sen Jia et al.ICLR 2021 · 86 citations
Builds on4
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras et al.ICCV 2019 · 668 citations
- Content and Style Disentanglement for Artistic Style TransferDmytro Kotovenko, Artsiom Sanakoyeu, Sabine Lang, Björn OmmerICCV 2019 · 187 citations
- Unsupervised Robust Disentangling of Latent Characteristics for Image SynthesisPatrick Esser, Johannes Haux, Björn OmmerICCV 2019 · 40 citations
- Interpreting the Latent Space of GANs for Semantic Face EditingYujun Shen, Jinjin Gu, Xiaoou Tang, Bolei ZhouCVPR 2020
Related papers
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao et al.AAAI 2022 · 19 citations
- Network-to-Network Translation with Conditional Invertible Neural NetworksRobin Rombach, Patrick Esser, Björn OmmerNeurIPS 2020 · 49 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Labeling Neural Representations with Inverse RecognitionKirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft et al.NeurIPS 2023 · 36 citations
- Explaining Deep Convolutional Neural Networks via Latent Visual-Semantic Filter AttentionYu Yang, Seungbae Kim, Jungseock JooCVPR 2022 · 11 citations
