GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S. Rodriguez, Jon E. Froehlich
Abstract
Voice assistants (VAs) like Siri and Alexa are transforming human-computer interaction; however, they lack awareness of users’ spatiotemporal context, resulting in limited performance and unnatural dialogue. We introduce GazePointAR, a fully-functional context-aware VA for wearable augmented reality that leverages eye gaze, pointing gestures, and conversation history to disambiguate speech queries. With GazePointAR, users can ask “what’s over there?” or “how do I solve this math problem?” simply by looking and/or pointing. We evaluated GazePointAR in a three-part lab study (N=12): (1) comparing GazePointAR to two commercial systems, (2) examining GazePointAR’s pronoun disambiguation across three tasks; (3) and an open-ended phase where participants could suggest and try their own context-sensitive queries. Participants appreciated the naturalness and human-like nature of pronoun-driven queries, although sometimes pronoun use was counter-intuitive. We then iterated on GazePointAR and conducted a first-person diary study examining how GazePointAR performs in-the-wild. We conclude by enumerating limitations and design considerations for future context-aware VAs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d02c3ed-97d1-4fbf-8d39-c1501735f5a6Cited by top-tier papers33
- Augmented Object Intelligence with XR-ObjectsMustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du et al.UIST 2024 · 64 citations
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou et al.CHI 2025 · 37 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart GlassesRunze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen et al.CHI 2025 · 26 citations
- OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question AnsweringJiahao Nick Li, Zhuohao Jerry Zhang, Jiaju MaCHI 2025 · 22 citations
Builds on12
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- Hand Interfaces: Using Hands to Imitate Objects in AR/VR for Expressive InteractionsSiyou Pei, Alexander Chen, Jaewook Lee, Yang ZhangCHI 2022 · 93 citations
- XAIR: A Framework of Explainable AI in Augmented RealityXuhai Xu, Anna Yu, Tanya R. Jonker, Kashyap Todi et al.CHI 2023 · 73 citations
- Evaluating the Potential of Glanceable AR Interfaces for Authentic Everyday UsesFeiyu Lu, Doug A. BowmanIEEE VR 2021 · 70 citations
Related papers
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented RealityYoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang et al.IEEE VR 2026 · 1 citation
- Towards More Transactional Voice Assistants: Investigating the Potential for a Multimodal Voice-Activated Indoor Navigation Assistant for Blind and Sighted TravelersAli Abdolrahmani, Maya Howes Gupta, Mei-Lian Vader, Ravi Kuber et al.CHI 2021 · 25 citations
- Enhancing Mobile Voice Assistants with WorldGazeSven Mayer, Gierad Laput, Chris HarrisonCHI 2020 · 65 citations
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao et al.UbiComp 2024 · 23 citations
- ARGESTUREAID: A Voice-Based, Adaptive, and Context-Aware Conversational Assistant for Supporting Mid-Air Gesture Discovery and ExecutionAnjali Khurana, Amy Karlson, Christopher Collins, Mengjie Yu et al.UbiComp 2026
