Voila-A: Aligning Vision-Language Models with User's Gaze Attention
Kun Yan, Zeyu Wang, Lei Ji, Yuntao Wang, Nan Duan, Shuai Ma
Abstract
In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling real-world applications with complex scenes and multiple objects, as well as aligning their focus with the diverse attention patterns of human users. In this paper, we introduce gaze information, feasibly collected by AR or VR devices, as a proxy for human attention to guide VLMs and propose a novel approach, Voila-A, for gaze alignment to enhance the interpretability and effectiveness of these models in real-world applications. First, we collect hundreds of minutes of gaze data to demonstrate that we can mimic human gaze modalities using localized narratives. We then design an automatic data annotation pipeline utilizing GPT-4 to generate the VOILA-COCO dataset. Additionally, we innovate the Voila Perceiver modules to integrate gaze information into VLMs while preserving their pretrained knowledge. We evaluate Voila-A using a hold-out validation set and a newly collected VOILA-GAZE Testset, which features real-life scenarios captured with a gaze-tracking device. Our experimental results demonstrate that Voila-A significantly outperforms several baseline models. By aligning model attention with human gaze patterns, Voila-A paves the way for more intuitive, user-centric VLMs and fosters engaging human-AI interaction across a wide range of applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao et al.UbiComp 2024 · 23 citations
- Gaze-VLM: Bridging Gaze and VLMs through Attention Regularization for Egocentric UnderstandingAnupam Pani, Yanchao YangNeurIPS 2025 · 12 citations
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin et al.ICCV 2025 · 8 citations
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang et al.CVPR 2026 · 8 citations
- AIDED: Augmenting Interior Design with Human Experience Data for Designer-AI Co-DesignYang Chen Lin, Chen-Ying Chien, Kai-Hsin Hou, Hung-Yu Chen et al.CHI 2026 · 2 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang et al.ICLR 2026 · 121 citations
- Empowering Large Language Models with 3D Situation AwarenessZhihao Yuan, Yibo Peng, Jinke Ren, Yinghong Liao et al.CVPR 2025
- VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical ReasoningNilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale et al.ICLR 2025
- Why are Visually-Grounded Language Models Bad at Image Classification?Yuhui Zhang, Alyssa Unell, Xiaohan Wang, Dhruba Ghosh et al.NeurIPS 2024 · 128 citations
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
