Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language
Zhenlin Xu, Marc Niethammer, Colin Raffel
Abstract
Deep learning models struggle with compositional generalization, i.e. the ability to recognize or generate novel combinations of observed elementary concepts. In hopes of enabling compositional generalization, various unsupervised learning algorithms have been proposed with inductive biases that aim to induce compositional structure in learned representations (e.g. disentangled representation and emergent language learning). In this work, we evaluate these unsupervised learning algorithms in terms of how well they enable compositional generalization. Specifically, our evaluation protocol focuses on whether or not it is easy to train a simple model on top of the learned representation that generalizes to new combinations of compositional factors. We systematically study three unsupervised representation learning algorithms -β-VAE, β-TCVAE, and emergent language (EL) autoencoders -on two datasets that allow directly testing compositional generalization. We find that directly using the bottleneck representation with simple models and few labels may lead to worse generalization than using representations from layers before or after the learned representation itself. In addition, we find that the previously proposed metrics for evaluating the levels of compositionality are not correlated with the actual compositional generalization in our framework. Surprisingly, we find that increasing pressure to produce a disentangled representation (e.g. increasing β in the β-VAE) produces representations with worse generalization, while representations from EL models show strong compositional generalization. Motivated by this observation, we further investigate the advantages of using EL to induce compositional structure in unsupervised representation learning, finding that it shows consistently stronger generalization than disentanglement models, especially when using less unlabeled data for unsupervised learning and fewer labels for downstream tasks. Taken together, our results shed new light onto the compositional generalization behavior of different unsupervised learning algorithms with a new setting to rigorously test this behavior, and suggest the potential benefits of developing EL learning algorithms for more generalizable representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1267e593-8723-4622-bbe7-5dbbfff93c09Cited by top-tier papers22
- Discovering modular solutions that generalize compositionallySimon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wolczyk et al.ICLR 2024 · 24 citations
- Improving Compositional Generalization using Iterated Learning and Simplicial EmbeddingsYi Ren, Samuel Lavoie, Michael Galkin, Danica J. Sutherland et al.NeurIPS 2023 · 21 citations
- Can Models Learn Skill Composition from Examples?Haoyu Zhao, Simran Kaur, Dingli Yu, Anirudh Goyal et al.NeurIPS 2024 · 20 citations
- How Diffusion Models Learn to Factorize and ComposeQiyao Liang, Ziming Liu, Mitchell Ostrow, Ila FieteNeurIPS 2024 · 17 citations
- Lewis's Signaling Game as beta-VAE For Natural Word Lengths and SegmentsRyo Ueda, Tadahiro TaniguchiICLR 2024 · 13 citations
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- Learning Compositional Representations for Few-Shot RecognitionPavel Tokmakov, Yu-Xiong Wang, Martial HebertICCV 2019 · 133 citations
Related papers
- Compositionality and Generalization In Emergent LanguagesRahma Chaabouni, Eugene Kharitonov, Diane Bouchacourt, Emmanuel Dupoux et al.ACL 2020 · 40 citations
- The role of Disentanglement in GeneralisationMilton Llera Montero, Casimir J. H. Ludwig, Rui Ponte Costa, Gaurav Malhotra et al.ICLR 2021 · 97 citations
- Towards Building A Group-based Unsupervised Representation Disentanglement FrameworkTao Yang, Xuanchi Ren, Yuwang Wang, Wenjun Zeng et al.ICLR 2022 · 36 citations
- MAGANet: Achieving Combinatorial Generalization by Modeling a Group ActionGeonho Hwang, Jaewoong Choi, Hyunsoo Cho, Myungjoo KangICML 2023 · 4 citations
- When does compositional structure yield compositional generalization? A kernel theorySamuel Lippl, Kim StachenfeldICLR 2025
