DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs
Chengtai Li, Yuting He, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang
摘要
Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adapt-ability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To tackle the challenges, we propose a Dynamic and Scalable Reasoning Framework (DSRF) that greatly enhances the reasoning ability by widening the network instead of deepening it, and dynamically adjusting the reasoning network to better fit novel samples instead of a static network. Specifically, we design a Multi-View Reasoning Pyramid (MVRP) to capture complex rules through layered reasoning to focus features at each view on distinct combinations of attributes, widening the reasoning network to cover more attribute combinations analogous to complex reasoning rules. Additionally, we propose a Dynamic Domain-Contrast Prediction (DDCP) block to handle varying task-specific relationships dynamically by introducing a Gram matrix to model feature distributions, and a gate matrix to capture subtle domain differences between context and target features. Extensive experiments on six AVR tasks demonstrate DSRF’s superior performance, achieving state-of-the-art results under various settings. Code is available here: https://github.com/UNNCRoxLi/DSRF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 被引用 313 次
- Learning What Not to Segment: A New Perspective on Few-Shot SegmentationChunbo Lang, Gong Cheng, Binfei Tu, Junwei HanCVPR 2022 · 被引用 289 次
相关 Paper
- DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number ReasoningChengtai Li, Yee Yang Tan, Yuting He, Jianfeng Ren 等AAAI 2025 · 被引用 6 次
- Learning Visual Abstract Reasoning through Dual-Stream NetworksKai Zhao, Chang Xu, Bailu SiAAAI 2024 · 被引用 11 次
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 被引用 10 次
- One Self-Configurable Model to Solve Many Abstract Visual Reasoning ProblemsMikolaj Malkinski, Jacek MandziukAAAI 2024 · 被引用 10 次
