COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language Navigation
Siqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan, Qi Wu, Zhihua Wei, Jing Liu
摘要
Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional components such as external knowledge bases or map information to enhance performance. These additions, while boosting performance, also lead to larger models and increased computational costs. In this paper, to achieve both high performance and low computational costs, we propose a novel architecture with the COmbination of Selective MemOrization (COSMO). Specifically, COSMO integrates state-space modules and transformer modules, and incorporates two VLN-customized selective state space modules: the Round Selective Scan (RSS) and the Cross-modal Selective State Space Module (CS3). RSS facilitates comprehensive inter-modal interactions within a single scan, while the CS3 module adapts the selective state space module into a dual-stream architecture, thereby enhancing the acquisition of cross-modal interactions. Experimental validations on three mainstream VLN benchmarks, REVERIE, R2R, and R2R-CE, not only demonstrate competitive navigation performance of our model but also show a significant reduction in computational costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language NavigationShuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong 等ICLR 2026 · 被引用 124 次
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 被引用 13 次
它引用的顶会 Paper52
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language NavigationYu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang 等AAAI 2026
- Cross-modal Map Learning for Vision and Language NavigationGeorgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan 等CVPR 2022 · 被引用 2 次
- VLN BERT: A Recurrent Vision-and-Language BERT for NavigationYicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo 等CVPR 2021
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelChangxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu 等AAAI 2026 · 被引用 2 次
- Augmented Commonsense Knowledge for Remote Object GroundingBahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu 等AAAI 2024 · 被引用 21 次
