G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
Zeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao, Kun Yan, Yuhan Wang, Lei Ji, Xuhai Xu, Chun Yu
Abstract
Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze --- a modality deeply linked to user intent and increasingly accessible via gaze-tracking wearables --- remains underexplored. This paper introduces a novel gaze-facilitated information querying paradigm, named G-VOILA, which synergizes users' gaze, visual field, and voice-based natural language queries to facilitate a more intuitive querying process. In a user-enactment study involving 21 participants in 3 daily scenarios (p = 21, scene = 3), we revealed the ambiguity in users' query language and a gaze-voice coordination pattern in users' natural query behaviors with G-VOILA. Based on the quantitative and qualitative findings, we developed a design framework for the G-VOILA paradigm, which effectively integrates the gaze data with the in-situ querying context. Then we implemented a G-VOILA proof-of-concept using cutting-edge deep learning techniques. A follow-up user study (p = 16, scene = 2) demonstrates its effectiveness by achieving both higher objective score and subjective score, compared to a baseline without gaze data. We further conducted interviews and provided insights for future gaze-facilitated information querying systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e469dccf-aa08-4f0b-8674-15d8e8501f84Cited by top-tier papers6
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou et al.CHI 2025 · 37 citations
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart GlassesRunze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen et al.CHI 2025 · 26 citations
- Computing with Smart Rings: A Systematic Literature ReviewZeyu Wang, Ruotong Yu, Xiangyang Wang, Jiexin Ding et al.UbiComp 2025 · 14 citations
- Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI CollaborationLeixian Shen, Yifang Wang, Huamin Qu, Xing Xie et al.CHI 2026 · 3 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
Builds on26
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang et al.NeurIPS 2024 · 43 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader UsersChu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif et al.CHI 2026 · 1 citation
- Integrating Gaze and Speech for Enabling Implicit InteractionsAnam Ahmad Khan, Joshua Newn, James Bailey, Eduardo VellosoCHI 2022 · 16 citations
- Enhancing Mobile Voice Assistants with WorldGazeSven Mayer, Gierad Laput, Chris HarrisonCHI 2020 · 65 citations
