Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G. Shapiro, Ranjay Krishna
Abstract
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning about 2D object instances benefits from object detection. Yet, MLMs can not produce intermediate depth or boxes to reason over. Fine-tuning MLMs on relevant data doesn't generalize well and outsourcing computation to specialized vision tools is too compute-intensive and memoryinefficient. To address this, we introduce Perception Tokens, intrinsic image representations designed to assist reasoning tasks where language is insufficient. Perception tokens act as auxiliary reasoning tokens, akin to chain-ofthought prompts in language models. For example, in a depth-related task, an MLM augmented with perception tokens can reason by generating a depth map as tokens, enabling it to solve the problem effectively. We propose AU-RORA, a training method that augments MLMs with perception tokens for improved reasoning over visual inputs. AURORA leverages a VQVAE to transform intermediate image representations, such as depth maps into a tokenized format and bounding box tokens, which are then used in a multi-task training framework. AURORA achieves notable improvements across counting benchmarks: +10.8% on BLINK, +11.3% on CVBench, and +8.3% on SEED-Bench, outperforming fine-tuning approaches in generalization across datasets. It also improves on relative depth: over +6% on BLINK. With perception tokens, AURORA expands the scope of MLMs beyond language-based reasoning, paving the way for more effective visual reasoning capabilities. Code and data will be released at the project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db2d68ab-4c3e-4d82-ae8d-ef8a71f03c07Cited by top-tier papers24
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang et al.ICLR 2026 · 80 citations
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen et al.CVPR 2026 · 64 citations
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought ReasoningJiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li et al.ICLR 2026 · 51 citations
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen et al.CVPR 2026 · 30 citations
Builds on26
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Don’t Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception ProgramsMuhammad Kamran Janjua, Hugo Silva, Di Niu, Bahador RashidiCVPR 2026 · 2 citations
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao et al.ICLR 2026 · 7 citations
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo et al.ICLR 2026 · 45 citations
- Multi-modal Auto-regressive Modeling via Visual TokensTianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao et al.ACM MM 2024 · 1 citation
- VCoder: Versatile Vision Encoders for Multimodal Large Language ModelsJitesh Jain, Jianwei Yang, Humphrey ShiCVPR 2024
