Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning
Sungjune Kim, Gyeongrok Oh, Heeju Ko, Daehyun Ji, Dongwook Lee, Byung-Jun Lee, Sujin Jang, Sangpil Kim
摘要
Navigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for online VLN? In our view, effective adaptation requires three qualities: 1) flexibility in handling different navigation outcomes, 2) interactivity with external environment, and 3) maintaining a harmony between plasticity and stability. To address this, we introduce FEEDTTA, a novel TTA framework for online VLN utilizing feedback-based reinforcement learning. Specifically, FEEDTTA learns by maximizing binary episodic feedback, a practical setup in which the agent receives a binary scalar after each episode that indicates the success or failure of the navigation. Additionally, we propose a gradient regularization technique that leverages the binary structure of FEEDTTA to achieve a balance between plasticity and stability during adaptation. Our extensive experiments on challenging VLN benchmarks demonstrate the superior adaptability of FEEDTTA, even outperforming the stateof-the-art offline training methods in REVERIE benchmark with a single stream of learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Active Test-time Vision-Language NavigationHeeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon 等NeurIPS 2025 · 被引用 10 次
- All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker AdaptationXudong Wang, Gan Li, Zhiyu Liu, Yao Wang 等ICLR 2026 · 被引用 4 次
- ME: Continual Vision-and-Language Navigation via Mixture of Macro and Micro ExpertsYongliang Jiang, Huaidong Zhang, Xuandi Luo, Shengfeng HeICLR 2026
- Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action ModelsZehua Zang, Xi Wang, Fuchun Sun, Xiao Xu 等CVPR 2026
- Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language NavigationZixuan Hu, Xuantuo Huang, Yancheng Li, Yichun Hu 等ICML 2026
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- On-the-Fly VLA Adaptation via Test-Time Reinforcement LearningChangyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang 等ACL 2026 · 被引用 7 次
- Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationJunyu Gao, Xuan Yao, Changsheng XuICML 2024 · 被引用 22 次
- Realistic Test-Time Adaptation of Vision-Language ModelsMaxime Zanella, Clément Fuchs, Christophe De Vleeschouwer, Ismail Ben AyedCVPR 2025
- Test-Time Adaptation with Binary FeedbackTaeckyung Lee, Sorn Chottananurak, Junsu Kim, Jinwoo Shin 等ICML 2025
- Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsMatteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini 等NeurIPS 2024 · 被引用 47 次
