Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
Yubo Wang, Juntian Zhang, Yichen Wu, Yankai Lin, Nils Lukas, Yuhan Liu
摘要
While Chain-of-Thought empowers Large Vision-Language Models with multi-step reasoning, explicit textual rationales suffer from an information bandwidth bottleneck, where continuous visual details are discarded during discrete tokenization. Recent latent reasoning methods attempt to address this challenge, but often fall prey to premature semantic collapse due to rigid autoregressive objectives. In this paper, we propose Laser, a novel paradigm that reformulates visual deduction via Dynamic Windowed Alignment Learning(DWAL). Instead of forcing a point-wise prediction, Laser aligns the latent state with a dynamic validity window of future semantics. This mechanism enforces a "Forest-before-Trees" cognitive hierarchy, enabling the model to maintain a probabilistic superposition of global features before narrowing down to local details. Crucially, Laser maintains interpretability via decodable trajectories while stabilizing unconstrained learning via Self-Refined Superposition. Extensive experiments on 6 benchmarks demonstrate that Laser achieves state-of-the-art performance among latent reasoning methods, surpassing the strong baseline Monet by 5.03% on average. Notably, it achieves these gains with extreme efficiency, reducing inference tokens by more than 97%, while demonstrating robust generalization to out-of-distribution domains. We hope this work encourages a paradigm shift from explicit next-token prediction to latent visual reasoning. * Equal contribution. The order was decided by a coin flip.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong 等ICML 2026 · 被引用 27 次
- Geometric Decoupling: Diagnosing the Structural Instability of LatentYuanbang Liang, Zhengwen Chen, Yu-Kun LaiICML 2026
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- Internalizing Explicit Reasoning into Latent Space for Dense RetrievalJiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long 等SIGIR 2026
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho 等ICML 2026 · 被引用 7 次
- SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language ModelsYunlong Chu, Minglai Shao, Yuhang Liu, Bing Hao 等KDD 2026 · 被引用 2 次
- Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action ModelsDianqiao Lei, Lianlei ShanICML 2026
- LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action ModelZhuoyang Liu, Jiaming Liu, Hao Chen, Jiale Yu 等ICML 2026 · 被引用 27 次
