Compositional Law Parsing with Latent Random Functions
Fan Shi, Bin Li, Xiangyang Xue
摘要
Human cognition has compositionality. We understand a scene by decomposing the scene into different concepts (e.g., shape and position of an object) and learning the respective laws of these concepts, which may be either natural (e.g., laws of motion) or man-made (e.g., laws of a game). The automatic parsing of these laws indicates the model's ability to understand the scene, which makes law parsing play a central role in many visual tasks. This paper proposes a deep latent variable model for Compositional LAw Parsing (CLAP), which achieves the human-like compositionality ability through an encoding-decoding architecture to represent concepts in the scene as latent variables. CLAP employs concept-specific latent random functions instantiated with Neural Processes to capture the law of concepts. Our experimental results demonstrate that CLAP outperforms the baseline methods in multiple visual tasks such as intuitive physics, abstract visual reasoning, and scene representation. The law manipulation experiments illustrate CLAP's interpretability by modifying specific latent random functions on samples. For example, CLAP learns the laws of position-changing and appearance constancy from the moving balls in a scene, making it possible to exchange laws between samples or compose existing laws into novel laws.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionFan Shi, Bin Li, Xiangyang XueICLR 2024 · 被引用 5 次
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu 等ICLR 2026
它引用的顶会 Paper12
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun 等ICLR 2020 · 被引用 276 次
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle 等NeurIPS 2020 · 被引用 167 次
- Improving Generative Imagination in Object-Centric World ModelsZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Bofeng Fu 等ICML 2020 · 被引用 97 次
- Meta-Learning Stationary Stochastic Process Prediction with Convolutional Neural ProcessesAndrew Y. K. Foong, Wessel P. Bruinsma, Jonathan Gordon, Yann Dubois 等NeurIPS 2020 · 被引用 96 次
相关 Paper
- Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2022 · 被引用 12 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi 等NeurIPS 2023 · 被引用 26 次
- Raven's Progressive Matrices Completion with Latent Gaussian Process PriorsFan Shi, Bin Li, Xiangyang XueAAAI 2021 · 被引用 10 次
- CtD: Composition through Decomposition in Emergent CommunicationBoaz Carmeli, Ron Meir, Yonatan BelinkovICLR 2025
