HOLM: Hallucinating Objects with Language Models for Referring Expression Recognition in Partially-Observed Scenes
Volkan Cirik, Louis-Philippe Morency, Taylor Berg-Kirkpatrick
摘要
AI systems embodied in the physical world face a fundamental challenge of partial observability; operating with only a limited view and knowledge of the environment. This creates challenges when AI systems try to reason about language and its relationship with the environment: objects referred to through language (e.g. giving many instructions) are not immediately visible. Actions by the AI system may be required to bring these objects in view. A good benchmark to study this challenge is Dynamic Referring Expression Recognition (dRER) task, where the goal is to find a target location by dynamically adjusting the field of view (FoV) in a partially observed 360 scenes. In this paper, we introduce HOLM, Hallucinating Objects with Language Models, to address the challenge of partial observability. HOLM uses large pre-trained language models (LMs) to infer object hallucinations for the unobserved part of the environment. Our core intuition is that if a pair of objects co-appear in an environment frequently, our usage of language should reflect this fact about the world. Based on this intuition, we prompt language models to extract knowledge about object affinities which gives us a proxy for spatial relationships of objects. Our experiments show that HOLM performs better than the state-of-the-art approaches on two datasets for dRER; allowing to study generalization for both indoor and outdoor settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Layout-Aware Dreamer for Embodied Visual Referring Expression GroundingMingxiao Li, Zehao Wang, Tinne Tuytelaars, Marie-Francine MoensAAAI 2023 · 被引用 22 次
- Believing is Seeing: Unobserved Object Detection using Generative ModelsSubhransu S. Bhattacharjee, Dylan Campbell, Rahul ShomeCVPR 2025
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
相关 Paper
- Refer360: A Referring Expression Recognition Dataset in 360 ImagesVolkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe MorencyACL 2020 · 被引用 11 次
- ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained EnvironmentsDong Wang, Xinghang Li, Zhengshen Zhang, Jirong Liu 等EMNLP 2025
- LVLMs are Bad at Overhearing Human Referential CommunicationZhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow 等EMNLP 2025
- Holodeck: Language Guided Generation of 3D Embodied AI EnvironmentsYue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt 等CVPR 2024 · 被引用 47 次
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 被引用 78 次
