How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot Interaction
Lesong Jia, Makayla Chang, Yu Liu, Na Du
摘要
Current multimodal instruction-recognition algorithms in human-robot interaction, developed largely from a purely technical perspective, remain rigid and incomplete in their use of human communicative cues. Therefore, a full understanding of how humans naturally refer to targets in interaction is central to enabling robots to interpret and act on user instructions. To investigate this, we collected multimodal behavior data from 30 participants who naturally instructed a robot for household tasks while we systematically varied target distance, direction, and local referent complexity. Our results show that speech instructions were often vague and lacked explicit target-position information. To resolve this ambiguity, multimodal cues are essential: gaze direction provides an order-of-magnitude improvement in target-localization accuracy, while hand pointing, head turns, and speech onset offer reliable temporal anchors for identifying target-directed gaze. We also found that speech patterns varied with distance and local referent complexity, whereas multimodal behaviors shifted with target direction, underscoring the need for context-adaptive recognition and interface design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Gaze-Supported 3D Object Manipulation in Virtual RealityDifeng Yu, Xueshi Lu, Rongkai Shi, Hai-Ning Liang 等CHI 2021 · 被引用 116 次
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 等ICCV 2021 · 被引用 57 次
- Robotic Characterization of Markerless Hand-Tracking on Meta Quest Pro and Quest 3 Virtual Reality HeadsetsEric Godden, William Steedman, Matthew K. X. J. PanIEEE VR 2025 · 被引用 20 次
- Integrating Gaze and Speech for Enabling Implicit InteractionsAnam Ahmad Khan, Joshua Newn, James Bailey, Eduardo VellosoCHI 2022 · 被引用 16 次
相关 Paper
- That and There: Judging the Intent of Pointing Actions with Robotic ArmsMalihe Alikhani, Baber Khalid, Rahul Shome, Chaitanya Mitash 等AAAI 2020
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou 等CHI 2026 · 被引用 1 次
- Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsKaran Ahuja, Andy Kong, Mayank Goel, Chris HarrisonUIST 2020 · 被引用 29 次
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented RealityYoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. KaufmanIEEE VR 2026
- Refer360: A Referring Expression Recognition Dataset in 360 ImagesVolkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe MorencyACL 2020 · 被引用 11 次
