Gaze-VLM: Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
Anupam Pani, Yanchao Yang
Abstract
Eye gaze offers valuable cues about attention, short-term intent, and future actions, making it a powerful signal for modeling egocentric behavior. In this work, we propose a gaze-regularized framework that enhances VLMs for two key egocentric understanding tasks: fine-grained future event prediction and current activity understanding. Unlike prior approaches that rely solely on visual inputs or use gaze as an auxiliary input signal , our method uses gaze only during training. We introduce a gaze-regularized attention mechanism that aligns model focus with human visual gaze. This design is flexible and modular, allowing it to generalize across multiple VLM architectures that utilize attention. Experimental results show that our approach improves semantic prediction scores by up to 11 for future event prediction and around 7 for current activity understanding, compared to the corresponding baseline models trained without gaze regularization. These results highlight the value of gaze-guided training in improving the accuracy and robustness of egocentric VLMs. Overall, this work establishes a foundation for using human gaze to enhance the predictive capabilities of VLMs in real-world scenarios like assistive robots and human-machine collaboration. Code and additional information is available at: https://github.com/anupampani/Gaze-VLM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 975da41a-bf47-4632-9323-5ba687bcc711Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
Related papers
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang et al.NeurIPS 2024 · 43 citations
- Eye Tracking-based LSTM for Locomotion Prediction in VRNiklas Stein, Gianni Bremer, Markus LappeIEEE VR 2022 · 39 citations
- ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic ManipulationZhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue et al.CVPR 2026 · 23 citations
- Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human ActivitiesMichele Mazzamuto, Antonino Furnari, Yoichi Sato, Giovanni Maria FarinellaCVPR 2025
- ActiveEye: Enabling Continuous and Responsive Video Understanding for Smart Eyewear SystemsZhenyu Xu, Tianlin Lu, Yingying Zhao, Yujiang Wang et al.UbiComp 2026 · 1 citation
