Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation
Bingqian Lin, Yi Zhu, Xiaodan Liang, Liang Lin, Jianzhuang Liu
摘要
Vision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existing VLN agents directly learn to align the raw directional features and visual features trained using one-hot labels to linguistic instruction features. However, the big semantic gap among these multi-modal inputs makes the alignment difficult and therefore limits the navigation performance. In this paper, we propose Actional Atomic-Concept Learning (AACL), which maps visual observations to actional atomic concepts for facilitating the alignment. Specifically, an actional atomic concept is a natural language phrase containing an atomic action and an object, e.g., ``go up stairs''. These actional atomic concepts, which serve as the bridge between observations and instructions, can effectively mitigate the semantic gap and simplify the alignment. AACL contains three core components: 1) a concept mapping module to map the observations to the actional atomic concept representations through the VLN environment and the recently proposed Contrastive Language-Image Pretraining (CLIP) model, 2) a concept refining adapter to encourage more instruction-oriented object concept extraction by re-ranking the predicted object concepts by CLIP, and 3) an observation co-embedding module which utilizes concept representations to regularize the observation representations. Our AACL establishes new state-of-the-art results on both fine-grained (R2R) and high-level (REVERIE and R2R-Last) VLN benchmarks. Moreover, the visualization shows that AACL significantly improves the interpretability in action decision. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/VLN-AACL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language NavigationSiqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan 等ICCV 2025 · 被引用 3 次
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language NavigationPeiran Xu, Xicheng Gong, Yadong MuICCV 2025 · 被引用 2 次
- Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language NavigationXuesong Zhang, Yunbo Xu, Jia Li, Ruonan Liu 等AAAI 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
相关 Paper
- ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsBingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang 等CVPR 2022 · 被引用 45 次
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 被引用 3 次
- Contrastive Instruction-Trajectory Learning for Vision-Language NavigationXiwen Liang, Fengda Zhu, Yi Zhu, Bingqian Lin 等AAAI 2022 · 被引用 29 次
- RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive LearningYi Zhu, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang 等EMNLP 2022 · 被引用 8 次
- Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional UnderstandingLe Zhang, Rabiul Awal, Aishwarya AgrawalCVPR 2024 · 被引用 7 次
