CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
Yiqi Zhu, Ziyue Wang, Can Zhang, Peng Li, Yang Liu
2025Year
Abstract
Figure 1. Illustration of continuous space perception: The left side depicts the construction of an image sequence representing a continuous space. The right side presents an example task related to the continuous space, showing how adjacent images are connected. Note that to answer the question correctly, one must recognize that the cars in images 1 and 2 are the same, as well as those in images 3 and 4.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a26b0f8-d79d-48ec-b0f0-1724771d31a4Builds on22
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningHaozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma et al.ICLR 2024 · 206 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao et al.ICLR 2024 · 95 citations
Related papers
- Continuous 3D Perception Model with Persistent StateQianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros et al.CVPR 2025
- Developers' Visuo-spatial Mental Model and Program ComprehensionAbir Bouraffa, Gian-Luca Fuhrmann, Walid MaalejICSE 2023 · 1 citation
- Continuous Scene Representations for Embodied AISamir Yitzhak Gadre, Kiana Ehsani, Shuran Song, Roozbeh MottaghiCVPR 2022 · 40 citations
- SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal ReasoningXiaoda Yang, Shenzhou Gao, Can Wang, Jiahe Zhang et al.AAAI 2026
- Can LLMs Learn to Map the World from Local Descriptions?Sirui Xia, Aili Chen, Xintao Wang, Tinghui Zhu et al.ACL 2026 · 2 citations
