Learning Visual Abstract Reasoning through Dual-Stream Networks
Kai Zhao, Chang Xu, Bailu Si
Abstract
Visual abstract reasoning tasks present challenges for deep neural networks, exposing limitations in their capabilities. In this work, we present a neural network model that addresses the challenges posed by Raven’s Progressive Matrices (RPM). Inspired by the two-stream hypothesis of visual processing, we introduce the Dual-stream Reasoning Network (DRNet), which utilizes two parallel branches to capture image features. On top of the two streams, a reasoning module first learns to merge the high-level features of the same image. Then, it employs a rule extractor to handle combinations involving the eight context images and each candidate image, extracting discrete abstract rules and utilizing an multilayer perceptron (MLP) to make predictions. Empirical results demonstrate that the proposed DRNet achieves state-of-the-art average performance across multiple RPM benchmarks. Furthermore, DRNet demonstrates robust generalization capabilities, even extending to various out-of-distribution scenarios. The dual streams within DRNet serve distinct functions by addressing local or spatial information. They are then integrated into the reasoning module, leveraging abstract rules to facilitate the execution of visual reasoning tasks. These findings indicate that the dual-stream architecture could play a crucial role in visual abstract reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd7147ba-fc1f-472a-80b9-51d4cb14194eCited by top-tier papers3
- DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMsChengtai Li, Yuting He, Jianfeng Ren, Ruibin Bai et al.NeurIPS 2025 · 2 citations
- DualMPNN: Harnessing Structural Alignments for High-Recovery Inverse Protein FoldingXuhui Liao, Qiyu Wang, Zhiqiang Liang, Liwei Xiao et al.NeurIPS 2025 · 2 citations
- GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEsKalliopi Basioti, Pritish Sahu, Tony Qingze Liu, Zihao Xu et al.ICLR 2025
Builds on14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao et al.NeurIPS 2023 · 810 citations
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei et al.AAAI 2021 · 126 citations
Related papers
- Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence TestsLingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang et al.ICML 2023 · 16 citations
- Effective Abstract Reasoning with Dual-Contrast NetworkTao Zhuo, Mohan S. KankanhalliICLR 2021 · 48 citations
- Learning to reason over visual objectsShanka Subhra Mondal, Taylor Whittington Webb, Jonathan CohenICLR 2023 · 7 citations
- Cognitive Predictive Coding Network: Rethinking the Generalization in Raven's Progressive MatricesXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang et al.ACM MM 2025
- Abstract Diagrammatic Reasoning with Multiplex Graph NetworksDuo Wang, Mateja Jamnik, Pietro LiòICLR 2020 · 74 citations
