Compositional Law Parsing with Latent Random Functions
Fan Shi, Bin Li, Xiangyang Xue
Abstract
Human cognition has compositionality. We understand a scene by decomposing the scene into different concepts (e.g., shape and position of an object) and learning the respective laws of these concepts, which may be either natural (e.g., laws of motion) or man-made (e.g., laws of a game). The automatic parsing of these laws indicates the model's ability to understand the scene, which makes law parsing play a central role in many visual tasks. This paper proposes a deep latent variable model for Compositional LAw Parsing (CLAP), which achieves the human-like compositionality ability through an encoding-decoding architecture to represent concepts in the scene as latent variables. CLAP employs concept-specific latent random functions instantiated with Neural Processes to capture the law of concepts. Our experimental results demonstrate that CLAP outperforms the baseline methods in multiple visual tasks such as intuitive physics, abstract visual reasoning, and scene representation. The law manipulation experiments illustrate CLAP's interpretability by modifying specific latent random functions on samples. For example, CLAP learns the laws of position-changing and appearance constancy from the moving balls in a scene, making it possible to exchange laws between samples or compose existing laws into novel laws.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31471c1b-0a8d-45ea-b988-03fccac5446fCited by top-tier papers3
- Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionFan Shi, Bin Li, Xiangyang XueICLR 2024 · 5 citations
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu et al.ICLR 2026
Builds on12
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle et al.NeurIPS 2020 · 167 citations
- Improving Generative Imagination in Object-Centric World ModelsZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Bofeng Fu et al.ICML 2020 · 97 citations
- Meta-Learning Stationary Stochastic Process Prediction with Convolutional Neural ProcessesAndrew Y. K. Foong, Wessel P. Bruinsma, Jonathan Gordon, Yann Dubois et al.NeurIPS 2020 · 96 citations
Related papers
- Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2022 · 12 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi et al.NeurIPS 2023 · 26 citations
- Raven's Progressive Matrices Completion with Latent Gaussian Process PriorsFan Shi, Bin Li, Xiangyang XueAAAI 2021 · 10 citations
- CtD: Composition through Decomposition in Emergent CommunicationBoaz Carmeli, Ron Meir, Yonatan BelinkovICLR 2025
