Abstract Spatial-Temporal Reasoning via Probabilistic Abduction and Execution
Chi Zhang, Baoxiong Jia, Song-Chun Zhu, Yixin Zhu
Abstract
Spatial-temporal reasoning is a challenging task in Artificial Intelligence (AI) due to its demanding but unique nature: a theoretic requirement on representing and reasoning based on spatial-temporal knowledge in mind, and an applied requirement on a high-level cognitive system capable of navigating and acting in space and time. Recent works have focused on an abstract reasoning task of this kind-Raven's Progressive Matrices (RPM). Despite the encouraging progress on RPM that achieves human-level performance in terms of accuracy, modern approaches have neither a treatment of human-like reasoning on generalization, nor a potential to generate answers. To fill in this gap, we propose a neuro-symbolic Probabilistic Abduction and Execution (PrAE) learner; central to the PrAE learner is the process of probabilistic abduction and execution on a probabilistic scene representation, akin to the mental manipulation of objects. Specifically, we disentangle perception and reasoning from a monolithic model. The neural visual perception frontend predicts objects' attributes, later aggregated by a scene inference engine to produce a probabilistic scene representation. In the symbolic logical reasoning backend, the PrAE learner uses the representation to abduce the hidden rules. An answer is predicted by executing the rules on the probabilistic representation. The entire system is trained end-to-end in an analysis-by-synthesis manner without any visual attribute annotations. Extensive experiments demonstrate that the PrAE learner improves cross-configuration generalization and is capable of rendering an answer, in contrast to prior works that merely make a categorical choice from candidates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd8eb422-cb03-4eb9-9b9d-c220a032d62aCited by top-tier papers18
- MEWL: Few-shot multimodal word learning with referential uncertaintyGuangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang et al.ICML 2023 · 29 citations
- GENOME: Generative Neuro-Symbolic Visual Reasoning by Growing and Reusing ModulesZhenfang Chen, Rui Sun, Wenjun Liu, Yining Hong et al.ICLR 2024 · 24 citations
- Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesZhuowan Li, Elias Stengel-Eskin, Yixiao Zhang, Cihang Xie et al.ICCV 2021 · 19 citations
- Active Reasoning in an Open-World EnvironmentManjie Xu, Guangyuan Jiang, Wei Liang, Chi Zhang et al.NeurIPS 2023 · 18 citations
- Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence TestsLingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang et al.ICML 2023 · 16 citations
Builds on8
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu et al.ICCV 2019 · 130 citations
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei et al.AAAI 2021 · 126 citations
- Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic ReasoningQing Li, Siyuan Huang, Yining Hong, Yixin Chen et al.ICML 2020 · 93 citations
- Abstract Diagrammatic Reasoning with Multiplex Graph NetworksDuo Wang, Mateja Jamnik, Pietro LiòICLR 2020 · 74 citations
Related papers
- Learning to reason over visual objectsShanka Subhra Mondal, Taylor Whittington Webb, Jonathan CohenICLR 2023 · 7 citations
- GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEsKalliopi Basioti, Pritish Sahu, Tony Qingze Liu, Zihao Xu et al.ICLR 2025
- VAEL: Bridging Variational Autoencoders and Probabilistic Logic ProgrammingEleonora Misino, Giuseppe Marra, Emanuele SansoneNeurIPS 2022 · 38 citations
- Hierarchical Perceptual and Predictive Analogy-Inference Network for Abstract Visual ReasoningWentao He, Jianfeng Ren, Ruibin Bai, Xudong JiangACM MM 2024 · 4 citations
- Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionFan Shi, Bin Li, Xiangyang XueICLR 2024 · 5 citations
