Necessary Conditions for Compositional Generalization of Embedding Models
Arnas Uselis, Andrea Dittadi, Seong Joon Oh
摘要
Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Modern models are trained on massive datasets, yet these are vanishingly small compared to the full combinatorial space of possible data, raising the question of whether models can reliably generalize to unseen combinations. To formalize what this requires, we propose a set of practically motivated desiderata that any compositionally generalizing system must satisfy, and analyze their implications under standard training with linear classification heads. We show that these desiderata necessitate linear factorization, where representations decompose additively into per-concept components, and further imply near-orthogonality across factors. We establish dimension bounds that link the number of concepts to the geometry of representations. Empirically, we survey CLIP and SigLIP families, finding strong evidence for linear factorization, approximate orthogonality, and a tight correlation between the quality of factorization and compositional generalization. Together, our results identify the structural conditions that embeddings must satisfy for compositional generalization, and provide both theoretical clarity and empirical diagnostics for developing foundation models that generalize compositionally.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper42
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf 等ICML 2020 · 被引用 361 次
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner 等CVPR 2022 · 被引用 349 次
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang 等ICLR 2024 · 被引用 249 次
相关 Paper
- Does Data Scaling Lead to Visual Compositional Generalization?Arnas Uselis, Andrea Dittadi, Seong Joon OhICML 2025
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf 等NeurIPS 2024 · 被引用 37 次
- When and How Does CLIP Enable Domain and Compositional Generalization?Elias Kempf, Simon Schrodi, Max Argus, Thomas BroxICML 2025
- Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsMatthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille 等ICCV 2023 · 被引用 51 次
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 被引用 4 次
