YouRefIt: Embodied Reference Understanding with Language and Gesture
Yixin Chen, Qing Li, Deqian Kong, Yik Lun Kei, Song-Chun Zhu, Tao Gao, Yixin Zhu, Siyuan Huang
摘要
We study the machine’s understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with perspective-taking to identify which object is being referred to. To tackle this problem, we introduce YouRefIt, a new crowd-sourced dataset of embodied reference collected in various physical scenes; the dataset contains 4,195 unique reference clips in 432 indoor scenes. To the best of our knowledge, this is the first embodied reference dataset that allows us to study referring expressions in daily physical scenes to understand referential behavior, human communication, and human-robot interaction. We further devise two benchmarks for image-based and video-based embodied reference understanding. Comprehensive baselines and extensive experiments provide the very first result of machine perception on how the referring expressions and gestures affect the embodied reference understanding. Our results provide essential evidence that gestural cues are as critical as language cues in understanding the embodied reference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu 等NeurIPS 2022 · 被引用 207 次
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan 等CVPR 2024 · 被引用 42 次
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li 等CVPR 2024 · 被引用 38 次
- PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesMd Mofijul Islam, Alexi Gladstone, Tariq IqbalAAAI 2023 · 被引用 10 次
- DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using DeterminersClarence Lee, M. Ganesh Kumar, Cheston TanICCV 2023 · 被引用 3 次
它引用的顶会 Paper6
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- Graph-Structured Referring Expression Reasoning in the WildSibei Yang, Guanbin Li, Yizhou YuCVPR 2020
- Multi-Task Collaborative Network for Joint Referring Expression Comprehension and SegmentationGen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao 等CVPR 2020
- Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionZhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong 等CVPR 2020
相关 Paper
- Understanding Embodied Reference with Touch-Line TransformerYang Li, Xiaoxue Chen, Hao Zhao, Jiangtao Gong 等ICLR 2023 · 被引用 6 次
- Grounding Language in Multi-Perspective Referential CommunicationZineng Tang, Lingjun Mao, Alane SuhrEMNLP 2024 · 被引用 1 次
- Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingAtharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen 等CVPR 2025
- ScanERU: Interactive 3D Visual Grounding Based on Embodied Reference UnderstandingZiyang Lu, Yunqiang Pei, Guoqing Wang, Peiwei Li 等AAAI 2024 · 被引用 12 次
- RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4DShuhei Kurita, Naoki Katsura, Eri OnamiICCV 2023 · 被引用 26 次
