DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs
Chengtai Li, Yuting He, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang
Abstract
Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adapt-ability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To tackle the challenges, we propose a Dynamic and Scalable Reasoning Framework (DSRF) that greatly enhances the reasoning ability by widening the network instead of deepening it, and dynamically adjusting the reasoning network to better fit novel samples instead of a static network. Specifically, we design a Multi-View Reasoning Pyramid (MVRP) to capture complex rules through layered reasoning to focus features at each view on distinct combinations of attributes, widening the reasoning network to cover more attribute combinations analogous to complex reasoning rules. Additionally, we propose a Dynamic Domain-Contrast Prediction (DDCP) block to handle varying task-specific relationships dynamically by introducing a Gram matrix to model feature distributions, and a gate matrix to capture subtle domain differences between context and target features. Extensive experiments on six AVR tasks demonstrate DSRF’s superior performance, achieving state-of-the-art results under various settings. Code is available here: https://github.com/UNNCRoxLi/DSRF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39a42865-cd1d-4094-ba19-8f3826455ba0Builds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao et al.NeurIPS 2023 · 810 citations
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 457 citations
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 313 citations
- Learning What Not to Segment: A New Perspective on Few-Shot SegmentationChunbo Lang, Gong Cheng, Binfei Tu, Junwei HanCVPR 2022 · 289 citations
Related papers
- DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number ReasoningChengtai Li, Yee Yang Tan, Yuting He, Jianfeng Ren et al.AAAI 2025 · 6 citations
- Learning Visual Abstract Reasoning through Dual-Stream NetworksKai Zhao, Chang Xu, Bailu SiAAAI 2024 · 11 citations
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 10 citations
- One Self-Configurable Model to Solve Many Abstract Visual Reasoning ProblemsMikolaj Malkinski, Jacek MandziukAAAI 2024 · 10 citations
