How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
Zhen Yang, Ping Jian, Zhongbin Guo, Zuming Zhang, Chengzhi Li, Yonghong Deng, Xinyue Zhang, Wenpeng Lu
2026Year
Abstract
Table!! Unknown Figure 1 : For textual viewpoint rotation understanding, human can easily imagine the spatial scenario and achieve 100% accuracy, while LLMs and VLMs struggle to output the correct answer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4b59ac5-0b1e-4324-bf77-448d9de370e1Builds on24
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
Related papers
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall SpacesJihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han et al.CVPR 2025
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian et al.ICLR 2026 · 11 citations
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine DifferentiationArina Kharlamova, Bowei He, Chen Ma, Xue LiuICLR 2026 · 3 citations
- Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation DatasetQian Chen, Xianyin Zhang, Yanzhi Liu, Lifan Guo et al.ACL 2026
