Multi-Task Gaze Communication Understanding
Cheng Peng, Oya Çeliktutan
Abstract
Human gaze communication is complex, comprising atomic-level (e.g. mutual, share, etc.) and event-level (e.g. follow, aversion, etc.) behaviours. Various methods have been developed to analyse gaze communication in images, but they typically fall short of fully understanding the complexities of the human gaze in videos. In this paper, we present a multi-task, multimodal model based on Contrastive Language-Image Pre-training (CLIP), designed to jointly predict atomic-level and event-level gaze communication, along with gaze target estimation. Specifically, we leverage the Vision-Language model to capture and utilise the semantic information between the atomic-level and event-level gaze communication categories. Additionally, most datasets in this field lack comprehensive annotations for both levels of gaze communication and detailed gaze target information. Therefore, we present a fully annotated gaze communication dataset, GP-Static++. We validate our model on GP-Static++ and several publicly available datasets, demonstrating its state-of-the-art performance. The dataset and code are available at https://pengc98.github.io/Multi-Task-Gaze-Communication-Understanding/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bf096e2e-25c7-4d27-b57b-8ac7ad4f4578Related papers
- Understanding Human Gaze Communication by Spatio-Temporal Graph ReasoningLifeng Fan, Wenguan Wang, Song-Chun Zhu, Xinyu Tang et al.ICCV 2019 · 124 citations
- Differential Contrastive Training for Gaze EstimationLin Zhang, Yi Tian, Xiyun Wang, Wanru Xu et al.ACM MM 2025 · 5 citations
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard et al.NeurIPS 2024 · 24 citations
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li et al.ICML 2025
- Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts LearningQiang Wang, Junlong Du, Ke Yan, Shouhong DingACM MM 2023 · 26 citations
