Gaze Target Detection by Merging Human Attention and Activity Cues
Yaokun Yang, Yihan Yin, Feng Lu
Abstract
Despite achieving impressive performance, current methods for detecting gaze targets, which depend on visual saliency and spatial scene geometry, continue to face challenges when it comes to detecting gaze targets within intricate image backgrounds. One of the primary reasons for this lies in the oversight of the intricate connection between human attention and activity cues. In this study, we introduce an innovative approach that amalgamates the visual saliency detection with the body-part & object interaction both guided by the soft gaze attention. This fusion enables precise and dependable detection of gaze targets amidst intricate image backgrounds. Our approach attains state-of-the-art performance on both the Gazefollow benchmark and the GazeVideoAttn benchmark. In comparison to recent methods that rely on intricate 3D reconstruction of a single input image, our approach, which solely leverages 2D image information, still exhibits a substantial lead across all evaluation metrics, positioning it closer to human-level performance. These outcomes underscore the potent effectiveness of our proposed method in the gaze target detection task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 682ff088-a3d7-4fbc-9df5-6b151d7abf30Cited by top-tier papers3
- Multi-View Gaze Target EstimationQiaomu Miao, Vivek Raju Golani, Jingyi Xu, Progga Paromita Dutta et al.ICCV 2025 · 4 citations
- Gaze Target Estimation Anywhere with ConceptsXu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou et al.CVPR 2026 · 3 citations
- GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context EncodingYuki Kawana, Shintaro Shiba, Quan Kong, Norimasa KoboriCVPR 2025
Builds on9
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo et al.CVPR 2022 · 69 citations
- ESCNet: Gaze Target Detection with the Understanding of 3D ScenesJun Bao, Buyu Liu, Jun YuCVPR 2022 · 36 citations
Related papers
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen et al.CVPR 2021
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 38 citations
- Vergence Matching: Inferring Attention to Objects in 3D Environments for Gaze-Assisted SelectionLudwig Sidenmark, Christopher Clarke, Joshua Newn, Mathias N. Lystbæk et al.CHI 2023 · 29 citations
- Exploring Pose-Aware Human-Object Interaction via Hybrid LearningEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangCVPR 2024 · 10 citations
- Looking here or there? Gaze Following in 360-Degree ImagesYunhao Li, Wei Shen, Zhongpai Gao, Yucheng Zhu et al.ICCV 2021 · 24 citations
