FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze Estimation
Daosong Hu, Mingyue Cui, Kai Huang
摘要
Gaze direction serves as a pivotal indicator for assessing the level of driver attention. While image-based gaze estimation has been extensively researched, there has been a recent shift towards capturing gaze direction from video sequences. This approach encounters notable challenges, including the comprehension of the dynamic pupil evolution across frames and the extraction of head pose information from a relatively static background. To surmount these challenges, we introduce a dual-stream deep learning framework that explicitly models the displacement changes of the pupil through a fine-grained inter-frame attention mechanism and generates weights to adjust gaze embeddings. This technique transforms the face into a set of distinct patches and employs cross-attention to ascertain the correlation between pixel displacements in various patches and adjacent frames, thereby tracking spatial dynamics within the sequence. Our method is validated using two publicly available driver gaze datasets, and the results indicate that it achieves state-of-the-art performance or is on par with the best outcomes while reducing the parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- Generalizing Gaze Estimation with Rotation ConsistencyYiwei Bao, Yunfei Liu, Haofei Wang, Feng LuCVPR 2022 · 被引用 54 次
- GaTector: A Unified Framework for Gaze Object PredictionBinglu Wang, Tao Hu, Baoshan Li, Xiaojuan Chen 等CVPR 2022 · 被引用 3 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze EstimationYihua Cheng, Yaning Zhu, Zongji Wang, Hongquan Hao 等CVPR 2024
相关 Paper
- Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose ReconstructionRuixu Liu, Ju Shen, He Wang, Chen Chen 等CVPR 2020
- Unsupervised Gaze Representation Learning from Multi-view Face ImagesYiwei Bao, Feng LuCVPR 2024
- DVGaze: Dual-View Gaze EstimationYihua Cheng, Feng LuICCV 2023 · 被引用 28 次
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen 等CVPR 2021
- UVAGaze: Unsupervised 1-to-2 Views Adaptation for Gaze EstimationRuicong Liu, Feng LuAAAI 2024 · 被引用 8 次
