Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping Review
Anam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou, Andrea Bianchi, Eduardo Velloso, Hans Gellersen, Joshua Newn
摘要
Multimodal interaction has long promised to make interfaces more intuitive and effective by combining complementary inputs. Among these, gaze and speech form a compelling pairing: gaze provides rapid spatial grounding, while speech conveys rich semantic information. Together, they offer rich cues for understanding user behaviour and intent. Yet despite decades of exploration, the research remains fragmented, making this synthesis timely as these inputs mature and are integrated into consumer-ready devices. This scoping review examined 103 studies published between 1991 and 2025, organised into explicit, where users intentionally provide gaze and speech, and implicit, where systems leverage users' natural behaviours to support interaction. Across both, we identified recurring ways for combining gaze and speech to resolve ambiguity, ground references, and support adaptivity. We contribute a synthesis of research on their combined use while highlighting challenges of temporal alignment, fusion and privacy, offering guidance for future research toward richer multimodal human-computer interaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Enhancing Mobile Voice Assistants with WorldGazeSven Mayer, Gierad Laput, Chris HarrisonCHI 2020 · 被引用 65 次
- Hummer: Text Entry by Gaze and HumRamin Hedeshy, Chandan Kumar, Raphael Menges, Steffen StaabCHI 2021 · 被引用 25 次
- EmBARDiment: an Embodied AI Agent for Productivity in XRRiccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez 等IEEE VR 2025 · 被引用 19 次
- Leveraging Error Correction in Voice-based Text Entry by Talk-and-GazeKorok Sengupta, Sabin Bhattarai, Sayan Sarcar, I. Scott MacKenzie 等CHI 2020 · 被引用 14 次
相关 Paper
- Using Speech to Visualise Shared Gaze Cues in MR Remote CollaborationAllison Jing, Gun A. Lee, Mark BillinghurstIEEE VR 2022 · 被引用 21 次
- Integrating Gaze and Speech for Enabling Implicit InteractionsAnam Ahmad Khan, Joshua Newn, James Bailey, Eduardo VellosoCHI 2022 · 被引用 16 次
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao 等UbiComp 2024 · 被引用 23 次
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu 等CHI 2026 · 被引用 2 次
