Bird's-Eye-View Scene Graph for Vision-Language Navigation
Rui Liu, Xiaohan Wang, Wenguan Wang, Yi Yang
摘要
Vision-language navigation (VLN), which entails an agent to navigate 3D environments following human instructions, has shown great advances. However, current agents are built upon panoramic observations, which hinders their ability to perceive 3D scene geometry and easily leads to ambiguous selection of panoramic view. To address these limitations, we present a BEV Scene Graph (BSG), which leverages multi-step BEV representations to encode scene layouts and geometric cues of indoor environment under the supervision of 3D detection. During navigation, BSG builds a local BEV representation at each step and maintains a BEV-based global scene map, which stores and organizes all the online collected local BEV representations according to their topological relations. Based on BSG, the agent predicts a local BEV grid-level decision score and a global graph-level decision score, combined with a subview selection score on panoramic views, for more accurate action prediction. Our approach significantly outperforms state-of-the-art methods on REVERIE, R2R, and R4R, showing the potential of BEV perception in VLN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- Dreamwalker: Mental Planning for Continuous Vision-Language NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangICCV 2023 · 被引用 98 次
- IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionJunbo Yin, Jianbing Shen, Runnan Chen, Wei Li 等CVPR 2024 · 被引用 73 次
- DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang 等ICML 2024 · 被引用 70 次
- OctoNav: Towards Generalist Embodied NavigationChen Gao, Liankai Jin, Xingyu Peng, Jiazhao Zhang 等CVPR 2026 · 被引用 42 次
- Vision-Language Navigation with Energy-Based PolicyRui Liu, Wenguan Wang, Yi YangNeurIPS 2024 · 被引用 39 次
它引用的顶会 Paper53
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia 等NeurIPS 2022 · 被引用 762 次
相关 Paper
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 被引用 25 次
- 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language NavigationJianzhe Gao, Rui Liu, Wenguan WangICCV 2025 · 被引用 5 次
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 等NeurIPS 2020 · 被引用 167 次
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu 等CVPR 2026 · 被引用 10 次
- The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language NavigationYuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang 等ICCV 2021 · 被引用 87 次
