Learning to reason over visual objects
Shanka Subhra Mondal, Taylor Whittington Webb, Jonathan Cohen
摘要
A core component of human intelligence is the ability to identify abstract patterns inherent in complex, high-dimensional perceptual data, as exemplified by visual reasoning tasks such as Raven's Progressive Matrices (RPM). Motivated by the goal of designing AI systems with this capacity, recent work has focused on evaluating whether neural networks can learn to solve RPM-like problems. Previous work has generally found that strong performance on these problems requires the incorporation of inductive biases that are specific to the RPM problem format, raising the question of whether such models might be more broadly useful. Here, we investigated the extent to which a general-purpose mechanism for processing visual scenes in terms of objects might help promote abstract visual reasoning. We found that a simple model, consisting only of an object-centric encoder and a transformer reasoning module, achieved state-of-the-art results on both of two challenging RPM-like benchmarks (PGM and I-RAVEN), as well as a novel benchmark with greater visual complexity (CLEVR-Matrices). These results suggest that an inductive bias for object-centric processing may be a key component of abstract visual reasoning, obviating the need for problem-specific inductive biases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata 等NeurIPS 2024 · 被引用 101 次
- Neural-Logic Human-Object Interaction DetectionLiulei Li, Jianan Wei, Wenguan Wang, Yi YangNeurIPS 2023 · 被引用 54 次
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
- Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence TestsLingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang 等ICML 2023 · 被引用 16 次
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee 等ICLR 2024 · 被引用 15 次
它引用的顶会 Paper10
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- GENESIS-V2: Inferring Unordered Object Representations without Iterative RefinementMartin Engelcke, Oiwi Parker Jones, Ingmar PosnerNeurIPS 2021 · 被引用 143 次
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei 等AAAI 2021 · 被引用 126 次
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds 等NeurIPS 2021 · 被引用 87 次
- Abstract Diagrammatic Reasoning with Multiplex Graph NetworksDuo Wang, Mateja Jamnik, Pietro LiòICLR 2020 · 被引用 74 次
相关 Paper
- GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEsKalliopi Basioti, Pritish Sahu, Tony Qingze Liu, Zihao Xu 等ICLR 2025
- Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical ReasoningWentao He, Jialu Zhang, Jianfeng Ren, Ruibin Bai 等AAAI 2023 · 被引用 22 次
- Learning Visual Abstract Reasoning through Dual-Stream NetworksKai Zhao, Chang Xu, Bailu SiAAAI 2024 · 被引用 11 次
- Raven's Progressive Matrices Completion with Latent Gaussian Process PriorsFan Shi, Bin Li, Xiangyang XueAAAI 2021 · 被引用 10 次
- Cognitive Predictive Coding Network: Rethinking the Generalization in Raven's Progressive MatricesXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang 等ACM MM 2025
