Learning Visual Abstract Reasoning through Dual-Stream Networks
Kai Zhao, Chang Xu, Bailu Si
摘要
Visual abstract reasoning tasks present challenges for deep neural networks, exposing limitations in their capabilities. In this work, we present a neural network model that addresses the challenges posed by Raven’s Progressive Matrices (RPM). Inspired by the two-stream hypothesis of visual processing, we introduce the Dual-stream Reasoning Network (DRNet), which utilizes two parallel branches to capture image features. On top of the two streams, a reasoning module first learns to merge the high-level features of the same image. Then, it employs a rule extractor to handle combinations involving the eight context images and each candidate image, extracting discrete abstract rules and utilizing an multilayer perceptron (MLP) to make predictions. Empirical results demonstrate that the proposed DRNet achieves state-of-the-art average performance across multiple RPM benchmarks. Furthermore, DRNet demonstrates robust generalization capabilities, even extending to various out-of-distribution scenarios. The dual streams within DRNet serve distinct functions by addressing local or spatial information. They are then integrated into the reasoning module, leveraging abstract rules to facilitate the execution of visual reasoning tasks. These findings indicate that the dual-stream architecture could play a crucial role in visual abstract reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMsChengtai Li, Yuting He, Jianfeng Ren, Ruibin Bai 等NeurIPS 2025 · 被引用 2 次
- DualMPNN: Harnessing Structural Alignments for High-Recovery Inverse Protein FoldingXuhui Liao, Qiyu Wang, Zhiqiang Liang, Liwei Xiao 等NeurIPS 2025 · 被引用 2 次
- GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEsKalliopi Basioti, Pritish Sahu, Tony Qingze Liu, Zihao Xu 等ICLR 2025
它引用的顶会 Paper14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu 等NeurIPS 2022 · 被引用 288 次
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei 等AAAI 2021 · 被引用 126 次
相关 Paper
- Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence TestsLingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang 等ICML 2023 · 被引用 16 次
- Effective Abstract Reasoning with Dual-Contrast NetworkTao Zhuo, Mohan S. KankanhalliICLR 2021 · 被引用 48 次
- Learning to reason over visual objectsShanka Subhra Mondal, Taylor Whittington Webb, Jonathan CohenICLR 2023 · 被引用 7 次
- Cognitive Predictive Coding Network: Rethinking the Generalization in Raven's Progressive MatricesXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang 等ACM MM 2025
- Abstract Diagrammatic Reasoning with Multiplex Graph NetworksDuo Wang, Mateja Jamnik, Pietro LiòICLR 2020 · 被引用 74 次
