Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
Shuai Dong, Siyuan Wang, Xingyu Liu, Chenglin Li, Haowen Hou, Zhongyu Wei
Abstract
Interleaved reasoning paradigms enhance Multimodal Large Language Models (MLLMs) with visual feedback but are hindered by the prohibitive computational cost of re-encoding pixel-dense images. A promising alternative, latent visual reasoning, circumvents this bottleneck yet faces limitations: methods either fail to capture intermediate state evolution due to single-step, non-interleaved structures, or sacrifice precise perceptual modeling by over-compressing features. We introduce Interleaved Latent Visual Reasoning (ILVR), a framework that unifies dynamic state evolution with precise perceptual modeling. ILVR interleaves textual generation with latent visual representations that act as specific, evolving cues for subsequent reasoning. Specifically, we employ a self-supervision strategy where a momentum teacher model selectively distills relevant features from ground-truth intermediate images into sparse supervision targets. This adaptive selection mechanism guides the model to autonomously generate context-aware visual signals. Extensive experiments on multimodal reasoning benchmarks demonstrate that ILVR outperforms existing approaches, effectively bridging the gap between fine-grained perception and sequential multimodal reasoning. The code is available at https://github. com/XD111ds/ILVR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c402d9e-9ad0-4276-aab3-913b3633d3c1Cited by top-tier papers7
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong et al.ICML 2026 · 27 citations
- Forest Before Trees: Latent Superposition for Efficient Visual ReasoningYubo Wang, Juntian Zhang, Yichen Wu, Yankai Lin et al.ACL 2026 · 12 citations
- Show, Don't Tell: Morphing Latent Reasoning into Image GenerationHarold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang et al.ICML 2026 · 7 citations
- Imagination Helps Visual Reasoning, But Not Yet in Latent SpaceYou Li, Chi Chen, Yanghao Li, Fanhu Zeng et al.ICML 2026 · 6 citations
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei et al.AAAI 2025 · 36 citations
Related papers
- Thinking in Latent Space: Progressive Multimodal Simplification for Visual ReasoningYuesen Tang, Yiming Yang, Tengfei Bao, Yu TongICML 2026
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang et al.ICLR 2026 · 80 citations
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho et al.ICML 2026 · 7 citations
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsHengzhuang Li, Xinsong Zhang, QIMING PENG, Bin Luo et al.CVPR 2026 · 2 citations
- Latent Implicit Visual ReasoningKelvin Li, Chuyi Shang, Leonid Karlinsky, Rogério Feris et al.CVPR 2026 · 15 citations
