Situational Fusion of Visual Representation for Visual Navigation
William B. Shen, Danfei Xu, Yuke Zhu, Li Fei-Fei, Leonidas J. Guibas, Silvio Savarese
摘要
A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair'', the agent might need to identify a chair in a living room using semantics, follow along a hallway using vanishing point cues, and avoid obstacles using depth. Therefore, utilizing the appropriate visual perception abilities based on a situational understanding of the visual environment can empower these navigation models in unseen visual environments. We propose to train an agent to fuse a large set of visual representations that correspond to diverse visual perception abilities. To fully utilize each representation, we develop an action-level representation fusion scheme, which predicts an action candidate from each representation and adaptively consolidate these action candidates into the final action. Furthermore, we employ a data-driven inter-task affinity regularization to reduce redundancies and improve generalization. Our approach leads to a significantly improved performance in novel environments over ImageNet-pretrained baseline and other fusion methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Robust and Generalizable Visual Representation Learning via Random ConvolutionsZhenlin Xu, Deyi Liu, Junlin Yang, Colin Raffel 等ICLR 2021 · 被引用 268 次
- Bird's-Eye-View Scene Graph for Vision-Language NavigationRui Liu, Xiaohan Wang, Wenguan Wang, Yi YangICCV 2023 · 被引用 100 次
- Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual NavigationZiad Al-Halah, Santhosh K. Ramakrishnan, Kristen GraumanCVPR 2022 · 被引用 52 次
- Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationYi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang 等ICCV 2021 · 被引用 41 次
- Learning Active Camera for Multi-Object NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu 等NeurIPS 2022 · 被引用 40 次
相关 Paper
- Narrowing the Gap between Vision and Action in NavigationYue Zhang, Parisa KordjamshidiACM MM 2024 · 被引用 2 次
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 被引用 25 次
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu 等ICCV 2025 · 被引用 1 次
- SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsGengze Zhou, Yicong Hong, Zun Wang, Chongyang Zhao 等ICCV 2025 · 被引用 4 次
- Uncertainty-Aware Gaussian Map for Vision-Language NavigationJianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao 等ICLR 2026 · 被引用 3 次
