Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning
Yixian Shen, Zhiheng Yang, Qi Bi, Changshuo Wang, Shuai Wang, JIA-HONG HUANG, George Floros, Prayag Tiwari, Anuj Pathania
摘要
Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To address this challenge, we propose Spectral-Progressive Thought Flow (SpecFlow), a novel lightweight multimodal spatial reasoning framework that represents intermediate visual thoughts in a fixed-size discrete cosine space. By exploiting strong energy compaction, SpecFlow preserves global layout and relational structure while introducing highfrequency details only when increased spatial precision is required. To align visual state evolution with linguistic intent, classifier-free guidance enables autoregressive textual thoughts to steer flowbased updates of the visual workspace (state) without expanding the context. As a result, SpecFlow maintains a bounded visual workspace whose updates depend only on the current visual state and accumulated textual trace, enabling long-horizon inference with stable latency and memory usage independent of reasoning depth. Empirical results show that SpecFlow achieves competitive or superior reasoning performance while reducing computation and KV cache costs by up to 2.1×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
相关 Paper
- Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric RoutingYixian Shen, Qi Bi, Zihan Wang, Zhiheng Yang 等ICLR 2026
- WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local ViewsSeongMin Jin, Doo Seok JeongICML 2026 · 被引用 1 次
- MMSep: Efficient Multimodal Long-Generation Reasoning via Multimodal Separator CompressionMingjie Ma, Yichao Ma, Jiannan Cao, Changhong Li 等KDD 2026
- EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMsZhenghao Chen, Huiqun Wang, Di HuangCVPR 2026 · 被引用 4 次
- Spectral Heat Flow for Conservative Token Condensation in Vision-Language ModelsZhaoyang Li, Yanjun Li, Wangkai Li, Yujia Chen 等ICML 2026
