Token Warping Helps MLLMs Look from Nearby Viewpoints
Phillip Y. Lee, Chanho Park, Mingue Park, Seungwoo Yoo, Juil Koo, Minhyuk Sung
Abstract
Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from nearby viewpoints? While MLLMs perform well on a single image reasoning, they remain fragile to viewpoint changes because pixel-level warping is highly sensitive to small depth errors and often introduces geometric distortions. Drawing on theories of mental imagery that posit part-level structural representations as the basis for human perspective transformation, we examine whether image tokens in ViT-based MLLMs serve as an effective substrate for viewpoint warping. We compare two token-level transformation strategies, forward and backward warping, and find that backward token fetching, which selects tokens at target-view grid locations and retrieves their counterparts from the source view, achieves greater stability and better preserves semantic coherence under viewpoint shifts. Experiments on our proposed ViewBench benchmark demonstrate that token-level warping enables MLLMs to reason reliably from nearby viewpoints, while consistently outperforming all baselines, including pixel-warping approaches, MLLMs fine-tuned for spatial reasoning, and a generative warping method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea7202c-4f9f-415c-8a67-a5ee553f7dceBuilds on72
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Grounding Everything in Tokens for Multimodal Large Language ModelsXiangxuan Ren, Zhongdao Wang, Liping Hou, Pin Tang et al.CVPR 2026 · 2 citations
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu et al.ICCV 2025 · 1 citation
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery SimulationPhillip Y. Lee, Jihyeon Je, Chanho Park, Leonidas J. Guibas et al.ICCV 2025 · 6 citations
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng et al.AAAI 2026 · 35 citations
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng et al.ACL 2026 · 8 citations
