Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning
Yixian Shen, Zhiheng Yang, Qi Bi, Changshuo Wang, Shuai Wang, JIA-HONG HUANG, George Floros, Prayag Tiwari, Anuj Pathania
Abstract
Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To address this challenge, we propose Spectral-Progressive Thought Flow (SpecFlow), a novel lightweight multimodal spatial reasoning framework that represents intermediate visual thoughts in a fixed-size discrete cosine space. By exploiting strong energy compaction, SpecFlow preserves global layout and relational structure while introducing highfrequency details only when increased spatial precision is required. To align visual state evolution with linguistic intent, classifier-free guidance enables autoregressive textual thoughts to steer flowbased updates of the visual workspace (state) without expanding the context. As a result, SpecFlow maintains a bounded visual workspace whose updates depend only on the current visual state and accumulated textual trace, enabling long-horizon inference with stable latency and memory usage independent of reasoning depth. Empirical results show that SpecFlow achieves competitive or superior reasoning performance while reducing computation and KV cache costs by up to 2.1×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
Related papers
- Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric RoutingYixian Shen, Qi Bi, Zihan Wang, Zhiheng Yang et al.ICLR 2026
- WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local ViewsSeongMin Jin, Doo Seok JeongICML 2026 · 1 citation
- MMSep: Efficient Multimodal Long-Generation Reasoning via Multimodal Separator CompressionMingjie Ma, Yichao Ma, Jiannan Cao, Changhong Li et al.KDD 2026
- EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMsZhenghao Chen, Huiqun Wang, Di HuangCVPR 2026 · 4 citations
- Spectral Heat Flow for Conservative Token Condensation in Vision-Language ModelsZhaoyang Li, Yanjun Li, Wangkai Li, Yujia Chen et al.ICML 2026
