ESCNet: Gaze Target Detection with the Understanding of 3D Scenes
Jun Bao, Buyu Liu, Jun Yu
Abstract
This paper aims to address the single image gaze target detection problem. Conventional methods either focus on 2D visual cues or exploit additional depth information in a very coarse manner. In this work, we propose to explicitly and effectively model 3D geometry under challenging scenario where only 2D annotations are available. We first obtain 3D point clouds of given scene with estimated depth and reference objects. Then we figure out the front-most points in all possible 3D directions of given person. These points are later leveraged in our ESCNet model. Specifically, ESCNet consists of geometry and scene parsing modules. The former produces an initial heatmap inferring the probability that each front-most point has been looking at according to estimated 3D gaze direction. And the latter further explores scene contextual cues to regulate detection results. We validate our idea on two publicly available dataset, GazeFollow and VideoAttentionTarget, and demon-strate the state-of-the-art performance. Our method also beats the human in terms of AUC on GazeFollow. Our code can be found here https://github.com/bjj9/ESCNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d21d4fcc-24d7-4463-9be2-6681e3aaecbeCited by top-tier papers11
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 41 citations
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 38 citations
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard et al.NeurIPS 2024 · 24 citations
- Toward Semantic Gaze Target DetectionSamy Tafasca, Anshul Gupta, Victor Bros, Jean-Marc OdobezNeurIPS 2024 · 16 citations
- Sharingan: A Transformer Architecture for Multi-Person Gaze FollowingSamy Tafasca, Anshul Gupta, Jean-Marc OdobezCVPR 2024 · 15 citations
Builds on6
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based OptimizationCheng Zhang, Zhaopeng Cui, Cai Chen, Shuaicheng Liu et al.ICCV 2021 · 43 citations
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
- Peek-a-Boo: Occlusion Reasoning in Indoor Scenes With Plane RepresentationsZiyu Jiang, Buyu Liu, Samuel Schulter, Zhangyang Wang et al.CVPR 2020
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen et al.CVPR 2021
Related papers
- Gaze Target Detection by Merging Human Attention and Activity CuesYaokun Yang, Yihan Yin, Feng LuAAAI 2024 · 6 citations
- Looking here or there? Gaze Following in 360-Degree ImagesYunhao Li, Wei Shen, Zhongpai Gao, Yucheng Zhu et al.ICCV 2021 · 24 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo et al.CVPR 2022 · 69 citations
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu et al.CVPR 2023
