EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
Xiaochuan Liu, Xin Cheng, Yuchong Sun, Xiaoxue Wu, Ruihua Song, Hao Sun, Denghao Zhang
Abstract
Imitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building alive virtual characters. Previous studies aim to predict gaze trajectories when humans are free-viewing an image, searching for required targets, or looking for clues to answer questions in an image. While these tasks focus on visual-centric scenarios, humans move their gaze also along with audio signal inputs in more common scenarios. To fill this gap, we introduce a new task that predicts human gaze trajectories in a visual scene with synchronized audio inputs and provide a new dataset containing 20k gaze points from 8 subjects. To effectively integrate audio information and simulate the dynamic process of human gaze motion, we propose a novel learning framework called EyEar (Eye moving while Ear listening) based on physics-informed dynamics, which considers three key factors to predict gazes: eye inherent motion tendency, vision salient attraction, and audio semantic attraction. We also propose a probability density score to overcome the high individual variability of gaze trajectories, thereby improving the stabilization of optimization and the reliability of the evaluation. Experimental results show that EyEar outperforms all the baselines in the context of all evaluation metrics, thanks to the proposed components in the learning model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a48cafc-ca79-40fa-8364-2994a2983a6bCited by top-tier papers1
Ask how each one uses itBuilds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingShilong Liu, Shijia Huang, Feng Li, Hao Zhang et al.AAAI 2023 · 44 citations
- Grounded Language-Image Pre-trainingLiunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang et al.CVPR 2022
- Connecting What To Say With Where To Look by Modeling Human Attention TracesZihang Meng, Licheng Yu, Ning Zhang, Tamara L. Berg et al.CVPR 2021
Related papers
- S3: Speech, Script and Scene driven Head and Eye AnimationYifang Pan, Rishabh Agrawal, Karan SinghSIGGRAPH 2024 · 11 citations
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- Learning from Human Gaze: Human-like Robot Social Navigation in Dense CrowdsZhecheng Yu, Yan Lyu, Chen Yang, Tao Chen et al.AAAI 2026
- Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionYuchen Zhou, Linkai Liu, Chao GouCVPR 2024 · 13 citations
- Context-Aware Head-and-Eye Motion Generation with Diffusion ModelYuxin Shen, Manjie Xu, Wei LiangIEEE VR 2024 · 4 citations
