Where Are You? Localization from Embodied Dialog
Meera Hahn, Jacob Krantz, Dhruv Batra, Devi Parikh, James M. Rehg, Stefan Lee, Peter Anderson
摘要
We present WHERE ARE YOU? (WAY), a dataset of 6k dialogs in which two humans – an Observer and a Locator – complete a cooperative localization task. The Observer is spawned at random in a 3D environment and can navigate from first-person views while answering questions from the Locator. The Locator must localize the Observer in a detailed top-down map by asking questions and giving instructions. Based on this dataset, we define three challenging tasks: Localization from Embodied Dialog or LED (localizing the Observer from dialog history), Embodied Visual Dialog (modeling the Observer), and Cooperative Localization (modeling both agents). In this paper, we focus on the LED task – providing a strong baseline model with detailed ablations characterizing both dataset biases and the importance of various modeling choices. Our best model achieves 32.7% success at identifying the Observer's location within 3m in unseen buildings, vs. 70.4% for human Locators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Pathdreamer: A World Model for Indoor NavigationJing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge 等ICCV 2021 · 被引用 128 次
- Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at ScaleRam Ramrakhya, Eric Undersander, Dhruv Batra, Abhishek DasCVPR 2022 · 被引用 73 次
- GridToPix: Training Embodied Agents with Minimal SupervisionUnnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi 等ICCV 2021 · 被引用 25 次
- Episodic Memory Question AnsweringSamyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai 等CVPR 2022 · 被引用 23 次
- SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and CollaborationYan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper3
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas 等EMNLP 2020 · 被引用 74 次
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang 等CVPR 2020
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
相关 Paper
- DiaLoc: An Iterative Approach to Embodied Dialog LocalizationChao Zhang, Mohan Li, Ignas Budvytis, Stephan LiwickiCVPR 2024 · 被引用 2 次
- DialNav: Multi-Turn Dialog Navigation with a Remote GuideLeekyeung Han, Hyunji Min, Gyeom Hwangbo, Jonghyun Choi 等ICCV 2025 · 被引用 1 次
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 等ICCV 2021 · 被引用 57 次
- Grounding Language in Multi-Perspective Referential CommunicationZineng Tang, Lingjun Mao, Alane SuhrEMNLP 2024 · 被引用 1 次
- Finding Fallen Objects Via Asynchronous Audio-Visual IntegrationChuang Gan, Yi Gu, Siyuan Zhou, Jeremy Schwartz 等CVPR 2022 · 被引用 13 次
