Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing
Yixian Shen, Qi Bi, Zihan Wang, Zhiheng Yang, Changshuo Wang, Zhi Zhang, Prayag Tiwari, Andy D. Pimentel, Anuj Pathania
摘要
Recently, visualization-of-thought (VoT) has unlocked new opportunities for complex spatial reasoning in multimodal large language models (MLLMs) by complementing verbal reasoning with visual thinking. However, the autoregressive accumulation of lengthy and redundant tokens substantially increases computation and memory costs. In this paper, we present a new efficient framework for multimodal spatial reasoning, named DARE, designed to adaptively prune multimodal tokens across different network depths, reasoning hops, and modalities. First, DARE devises an intra- and inter-hop-aware differentiable retention mechanism to dynamically estimate token importance both within each reasoning step and across successive hops. Recognizing that deeper network layers encode visual cues into verbal streams, DARE introduces an asymmetric compression strategy that prunes tokens according to modality-specific redundancy and semantic importance. Furthermore, DARE incorporates a progressive KV-cache retention policy aligned with cross-modal fusion dynamics, further reducing memory overhead during autoregressive reasoning. Our method delivers substantial reductions in computation and memory footprint, averaging a 40.37% reduction in FLOPs and 46.07% reduction in KV caches usage, while consistently preserving or even improving reasoning performance across seven multimodal spatial reasoning benchmarks, and further generalizing to broader multimodal reasoning tasks, establishing a scalable and robust recipe for efficient multimodal reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsJitai Hao, Hao Liu, Xinyan Xiao, Qiang Huang 等ICLR 2026 · 被引用 18 次
- Spectral-Progressive Thought Flow for Lightweight Multimodal ReasoningYixian Shen, Zhiheng Yang, Qi Bi, Changshuo Wang 等ICML 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
相关 Paper
- PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal ModelsSong Li, yongping xiongICML 2026
- VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal ReasoningHengbo Xu, Shengjie Jin, Yanbiao Ma, Zhiwu LuICML 2026 · 被引用 1 次
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee 等ICCV 2025 · 被引用 37 次
- VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationSihan Yang, Runsen Xu, Chenhang Cui, Tai Wang 等ICCV 2025 · 被引用 1 次
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsLei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie 等EMNLP 2025 · 被引用 2 次
