Integrating Gaze and Speech for Enabling Implicit Interactions
Anam Ahmad Khan, Joshua Newn, James Bailey, Eduardo Velloso
Abstract
Gaze and speech are rich contextual sources of information that, when combined, can result in effective and rich multimodal interactions. This paper proposes a machine learning-based pipeline that leverages and combines users’ natural gaze activity, the semantic knowledge from their vocal utterances and the synchronicity between gaze and speech data to facilitate users’ interaction. We evaluated our proposed approach on an existing dataset, which involved 32 participants recording voice notes while reading an academic paper. Using a Logistic Regression classifier, we demonstrate that our proposed multimodal approach maps voice notes with accurate text passages with an average F1-Score of 0.90. Our proposed pipeline motivates the design of multimodal interfaces that combines natural gaze and speech patterns to enable robust interactions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d5972ff3-d0ee-413e-a241-b3936c9bb074Cited by top-tier papers5
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- PeerEdu: Bootstrapping Online Learning Behaviors via Asynchronous Area of Interest Sharing from Peer GazeSonglin Xu, Dongyin Hu, Ru Wang, Xinyu ZhangCHI 2025 · 9 citations
- PointAloud: An Interaction Suite for AI-Supported Pointer-Centric Think-Aloud ComputingFrederic Gmeiner, John Thompson, George W. Fitzmaurice, Justin MatejkaCHI 2026 · 1 citation
- Alfa: Attentive Low-Rank Filter Adaptation for Structure-Aware Cross-Domain Personalized Gaze EstimationHe-Yen Hsieh, Wei-Te Mark Ting, H. T. KungAAAI 2026 · 1 citation
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
Related papers
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou et al.CHI 2026 · 1 citation
- Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual SignalsYuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng ShiACL 2025 · 5 citations
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao et al.UbiComp 2024 · 23 citations
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang et al.NeurIPS 2024 · 43 citations
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
