FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze Estimation
Daosong Hu, Mingyue Cui, Kai Huang
Abstract
Gaze direction serves as a pivotal indicator for assessing the level of driver attention. While image-based gaze estimation has been extensively researched, there has been a recent shift towards capturing gaze direction from video sequences. This approach encounters notable challenges, including the comprehension of the dynamic pupil evolution across frames and the extraction of head pose information from a relatively static background. To surmount these challenges, we introduce a dual-stream deep learning framework that explicitly models the displacement changes of the pupil through a fine-grained inter-frame attention mechanism and generates weights to adjust gaze embeddings. This technique transforms the face into a set of distinct patches and employs cross-attention to ascertain the correlation between pixel displacements in various patches and adjacent frames, thereby tracking spatial dynamics within the sequence. Our method is validated using two publicly available driver gaze datasets, and the results indicate that it achieves state-of-the-art performance or is on par with the best outcomes while reducing the parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- Generalizing Gaze Estimation with Rotation ConsistencyYiwei Bao, Yunfei Liu, Haofei Wang, Feng LuCVPR 2022 · 54 citations
- GaTector: A Unified Framework for Gaze Object PredictionBinglu Wang, Tao Hu, Baoshan Li, Xiaojuan Chen et al.CVPR 2022 · 3 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze EstimationYihua Cheng, Yaning Zhu, Zongji Wang, Hongquan Hao et al.CVPR 2024
Related papers
- Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose ReconstructionRuixu Liu, Ju Shen, He Wang, Chen Chen et al.CVPR 2020
- Unsupervised Gaze Representation Learning from Multi-view Face ImagesYiwei Bao, Feng LuCVPR 2024
- DVGaze: Dual-View Gaze EstimationYihua Cheng, Feng LuICCV 2023 · 28 citations
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen et al.CVPR 2021
- UVAGaze: Unsupervised 1-to-2 Views Adaptation for Gaze EstimationRuicong Liu, Feng LuAAAI 2024 · 8 citations
