Lune

CHI2026顶会

How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot Interaction

Lesong Jia, Makayla Chang, Yu Liu, Na Du

2026年份

摘要

Current multimodal instruction-recognition algorithms in human-robot interaction, developed largely from a purely technical perspective, remain rigid and incomplete in their use of human communicative cues. Therefore, a full understanding of how humans naturally refer to targets in interaction is central to enabling robots to interpret and act on user instructions. To investigate this, we collected multimodal behavior data from 30 participants who naturally instructed a robot for household tasks while we systematically varied target distance, direction, and local referent complexity. Our results show that speech instructions were often vague and lacked explicit target-position information. To resolve this ambiguity, multimodal cues are essential: gaze direction provides an order-of-magnitude improvement in target-localization accuracy, while hand pointing, head turns, and speech onset offer reliable temporal anchors for identifying target-directed gaze. We also found that speech patterns varied with distance and local referent complexity, whereas multimodal behaviors shifted with target direction, underscoring the need for context-adaptive recognition and interface design.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖