Compositional Generalization via Forced Rendering of Disentangled Latents
Qiyao Liang, Daoyuan Qian, Liu Ziyin, Ila R. Fiete
摘要
Composition-the ability to generate myriad variations from finite means-is believed to underlie powerful generalization. However, compositional generalization remains a key challenge for deep learning. A widely held assumption is that learning disentangled (factorized) representations naturally supports this kind of extrapolation. Yet, empirical results are mixed, with many generative models failing to recognize and compose factors to generate out-of-distribution (OOD) samples. In this work, we investigate a controlled 2D Gaussian "bump" generation task with fully disentangled (x, y) inputs, demonstrating that standard generative architectures still fail in OOD regions when training with partial data, by reentangling latent representations in subsequent layers. By examining the model's learned kernels and manifold geometry, we show that this failure reflects a "memorization" strategy for generation via data superposition rather than via composition of the true factorized features. We show that when models are forced-through architectural modifications with regularization or curated training data-to render the disentangled latents into the full-dimensional representational (pixel) space, they can be highly data-efficient and effective at composing in OOD regions. These findings underscore that disentangled latents in an abstract representation are insufficient and show that if models can represent disentangled factors directly in the output representational space, it can achieve robust compositional generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
- Seeing Through the PRISM: Compound & Controllable Restoration of Scientific ImagesRupa Kurinchi-Vendhan, Pratyusha Sharma, Antonio Torralba, Sara BeeryICLR 2026
- Relational Structural Causal ModelsAdiba Ejaz, Elias BareinboimICML 2026
它引用的顶会 Paper8
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 被引用 113 次
- The role of Disentanglement in GeneralisationMilton Llera Montero, Casimir J. H. Ludwig, Rui Ponte Costa, Gaurav Malhotra 等ICLR 2021 · 被引用 97 次
- Compositional Generalization from First PrinciplesThaddäus Wiedemer, Prasanna Mayilvahanan, Matthias Bethge, Wieland BrendelNeurIPS 2023 · 被引用 78 次
- Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent LanguageZhenlin Xu, Marc Niethammer, Colin RaffelNeurIPS 2022 · 被引用 59 次
- Lost in Latent Space: Examining failures of disentangled models at combinatorial generalisationMilton Llera Montero, Jeffrey S. Bowers, Rui Ponte Costa, Casimir J. H. Ludwig 等NeurIPS 2022 · 被引用 29 次
相关 Paper
- When does compositional structure yield compositional generalization? A kernel theorySamuel Lippl, Kim StachenfeldICLR 2025
- How Diffusion Models Learn to Factorize and ComposeQiyao Liang, Ziming Liu, Mitchell Ostrow, Ila FieteNeurIPS 2024 · 被引用 17 次
- Is Generation Required for Data-Efficient Perception?Jack Brady, Bernhard Schölkopf, Thomas Kipf, Simon Buchholz 等ICML 2026 · 被引用 1 次
- Does Data Scaling Lead to Visual Compositional Generalization?Arnas Uselis, Andrea Dittadi, Seong Joon OhICML 2025
- MAGANet: Achieving Combinatorial Generalization by Modeling a Group ActionGeonho Hwang, Jaewoong Choi, Hyunsoo Cho, Myungjoo KangICML 2023 · 被引用 4 次
