Integrating Gaze and Speech for Enabling Implicit Interactions
Anam Ahmad Khan, Joshua Newn, James Bailey, Eduardo Velloso
摘要
Gaze and speech are rich contextual sources of information that, when combined, can result in effective and rich multimodal interactions. This paper proposes a machine learning-based pipeline that leverages and combines users’ natural gaze activity, the semantic knowledge from their vocal utterances and the synchronicity between gaze and speech data to facilitate users’ interaction. We evaluated our proposed approach on an existing dataset, which involved 32 participants recording voice notes while reading an academic paper. Using a Logistic Regression classifier, we demonstrate that our proposed multimodal approach maps voice notes with accurate text passages with an average F1-Score of 0.90. Our proposed pipeline motivates the design of multimodal interfaces that combines natural gaze and speech patterns to enable robust interactions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- PeerEdu: Bootstrapping Online Learning Behaviors via Asynchronous Area of Interest Sharing from Peer GazeSonglin Xu, Dongyin Hu, Ru Wang, Xinyu ZhangCHI 2025 · 被引用 9 次
- PointAloud: An Interaction Suite for AI-Supported Pointer-Centric Think-Aloud ComputingFrederic Gmeiner, John Thompson, George W. Fitzmaurice, Justin MatejkaCHI 2026 · 被引用 1 次
- Alfa: Attentive Low-Rank Filter Adaptation for Structure-Aware Cross-Domain Personalized Gaze EstimationHe-Yen Hsieh, Wei-Te Mark Ting, H. T. KungAAAI 2026 · 被引用 1 次
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
相关 Paper
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou 等CHI 2026 · 被引用 1 次
- Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual SignalsYuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng ShiACL 2025 · 被引用 5 次
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao 等UbiComp 2024 · 被引用 23 次
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang 等NeurIPS 2024 · 被引用 43 次
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
