See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
Pengteng Li, Pinhao Song, Wuyang Li, Huizai Yao, Weiyu Guo, Yijie Xu, Dugang Liu, Hui Xiong
Abstract
We introduce SEE&TREK, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMS) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visualspatial understanding remains underexplored. SEE&TREK addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPUfree, requiring only a single forward pass, and can be seamlessly integrated into existing MLLMS. Extensive experiments on the VSI-BENCH and STI-BENCH show that SEE&TREK consistently boosts various MLLMS performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence. The link of code: https://github.com/Hoantrbl/SeeTrek.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a17b0438-b05f-4657-816a-b6b0c823d933Cited by top-tier papers5
- EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMsZhenghao Chen, Huiqun Wang, Di HuangCVPR 2026 · 4 citations
- VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object DetectionWuyang Li, Zhu Yu, Alexandre AlahiNeurIPS 2025 · 3 citations
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual DistillationChiao-An Yang, Ryo Hachiuma, Sifei Liu, Subhashree Radhakrishnan et al.CVPR 2026 · 2 citations
- Reason, Then Re-reason: Cross-view Revisiting Improves Spatial ReasoningChaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng et al.ICML 2026 · 1 citation
- Token Warping Helps MLLMs Look from Nearby ViewpointsPhillip Y. Lee, Chanho Park, Mingue Park, Seungwoo Yoo et al.CVPR 2026 · 1 citation
Builds on26
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin et al.NeurIPS 2025 · 164 citations
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 135 citations
Related papers
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- Thinking with Geometry: Active Geometry Integration for Spatial ReasoningHaoyuan Li, Qihang Cao, Tao Tang, Kun Xiang et al.ICML 2026 · 12 citations
- Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided ReasoningJiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang et al.ACL 2026 · 5 citations
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen et al.CVPR 2026 · 64 citations
- Spatial Understanding from Videos: Structured Prompts Meet Simulation DataHaoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen et al.NeurIPS 2025 · 31 citations
