Layout-Aware Dreamer for Embodied Visual Referring Expression Grounding
Mingxiao Li, Zehao Wang, Tinne Tuytelaars, Marie-Francine Moens
Abstract
In this work, we study the problem of Embodied Referring Expression Grounding, where an agent needs to navigate in a previously unseen environment and localize a remote object described by a concise high-level natural language instruction. When facing such a situation, a human tends to imagine what the destination may look like and to explore the environment based on prior knowledge of the environmental layout, such as the fact that a bathroom is more likely to be found near a bedroom than a kitchen. We have designed an autonomous agent called Layout-aware Dreamer (LAD), including two novel modules, that is, the Layout Learner and the Goal Dreamer to mimic this cognitive decision process. The Layout Learner learns to infer the room category distribution of neighboring unexplored areas along the path for coarse layout estimation, which effectively introduces layout common sense of room-to-room transitions to our agent. To learn an effective exploration of the environment, the Goal Dreamer imagines the destination beforehand. Our agent achieves new state-of-the-art performance on the public leaderboard of REVERIE dataset in challenging unseen test environments with improvement on navigation success rate (SR) by 4.02% and remote grounding success (RGS) by 3.43% comparing to previous previous state of the art. The code is released at https://github.com/zehao-wang/LAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87b09adf-91b0-42d0-affc-a616f32131d1Cited by top-tier papers5
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai et al.ACL 2024 · 29 citations
- Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language NavigationYiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng WangAAAI 2025 · 8 citations
- COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language NavigationSiqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan et al.ICCV 2025 · 3 citations
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language NavigationPeiran Xu, Xicheng Gong, Yadong MuICCV 2025 · 2 citations
- Do Visual Imaginations Improve Vision-and-Language Navigation Agents?Akhil Perincherry, Jacob Krantz, Stefan LeeCVPR 2025
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- Scene-Intuitive Agent for Remote Embodied Visual GroundingXiangru Lin, Guanbin Li, Yizhou YuCVPR 2021
- Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring ExpressionChen Gao, Jinyu Chen, Si Liu, Luting Wang et al.CVPR 2021
- Augmented Commonsense Knowledge for Remote Object GroundingBahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu et al.AAAI 2024 · 21 citations
- Target-Driven Structured Transformer Planner for Vision-Language NavigationYusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang et al.ACM MM 2022 · 49 citations
- March in Chat: Interactive Prompting for Remote Embodied Referring ExpressionYanyuan Qiao, Yuankai Qi, Zheng Yu, Jing Liu et al.ICCV 2023 · 52 citations
