Toward a Visual Concept Vocabulary for GAN Latent Space
Sarah Schwettmann, Evan Hernandez, David Bau, Samuel Klein, Jacob Andreas, Antonio Torralba
Abstract
A large body of recent work has identified transformations in the latent spaces of generative adversarial networks (GANs) that consistently and interpretably transform generated images. But existing techniques for identifying these transformations rely on either a fixed vocabulary of prespecified visual concepts, or on unsupervised disentanglement techniques whose alignment with human judgments about perceptual salience is unknown. This paper introduces a new method for building open-ended vocabularies of primitive visual concepts represented in a GAN’s latent space. Our approach is built from three components: (1) automatic identification of perceptually salient directions based on their layer selectivity; (2) human annotation of these directions with free-form, compositional natural language descriptions; and (3) decomposition of these annotations into a visual concept vocabulary, consisting of distilled directions labeled with single words. Experiments show that concepts learned with our approach are reliable and composable—generalizing across classes, contexts, and observers, and enabling fine-grained manipulation of image style and content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Natural Language Descriptions of Deep Visual FeaturesEvan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili et al.ICLR 2022 · 160 citations
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like InteractionsJohn Joon Young Chung, Eytan AdarUIST 2023 · 94 citations
- Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMsQingru Zhang, Chandan Singh, Liyuan Liu, Xiaodong Liu et al.ICLR 2024 · 76 citations
- A Multimodal Automated Interpretability AgentTamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram et al.ICML 2024 · 57 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
- GANalyze: Toward Visual Definitions of Cognitive Image PropertiesLore Goetschalckx, Alex Andonian, Aude Oliva, Phillip IsolaICCV 2019 · 345 citations
Related papers
- Controlling generative models with continuous factors of variationsAntoine Plumerault, Hervé Le Borgne, Céline HudelotICLR 2020 · 132 citations
- GAN "Steerability" without optimizationNurit Spingarn, Ron Banner, Tomer MichaeliICLR 2021 · 59 citations
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao et al.AAAI 2022 · 19 citations
- Navigating the GAN Parameter Space for Semantic Image EditingAnton Cherepkov, Andrey Voynov, Artem BabenkoCVPR 2021
- LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable DirectionsOguz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er, Pinar YanardagICCV 2021 · 71 citations
