ESCNet: Gaze Target Detection with the Understanding of 3D Scenes
Jun Bao, Buyu Liu, Jun Yu
摘要
This paper aims to address the single image gaze target detection problem. Conventional methods either focus on 2D visual cues or exploit additional depth information in a very coarse manner. In this work, we propose to explicitly and effectively model 3D geometry under challenging scenario where only 2D annotations are available. We first obtain 3D point clouds of given scene with estimated depth and reference objects. Then we figure out the front-most points in all possible 3D directions of given person. These points are later leveraged in our ESCNet model. Specifically, ESCNet consists of geometry and scene parsing modules. The former produces an initial heatmap inferring the probability that each front-most point has been looking at according to estimated 3D gaze direction. And the latter further explores scene contextual cues to regulate detection results. We validate our idea on two publicly available dataset, GazeFollow and VideoAttentionTarget, and demon-strate the state-of-the-art performance. Our method also beats the human in terms of AUC on GazeFollow. Our code can be found here https://github.com/bjj9/ESCNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 被引用 41 次
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard 等NeurIPS 2024 · 被引用 24 次
- Toward Semantic Gaze Target DetectionSamy Tafasca, Anshul Gupta, Victor Bros, Jean-Marc OdobezNeurIPS 2024 · 被引用 16 次
- Sharingan: A Transformer Architecture for Multi-Person Gaze FollowingSamy Tafasca, Anshul Gupta, Jean-Marc OdobezCVPR 2024 · 被引用 15 次
它引用的顶会 Paper6
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based OptimizationCheng Zhang, Zhaopeng Cui, Cai Chen, Shuaicheng Liu 等ICCV 2021 · 被引用 43 次
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
- Peek-a-Boo: Occlusion Reasoning in Indoor Scenes With Plane RepresentationsZiyu Jiang, Buyu Liu, Samuel Schulter, Zhangyang Wang 等CVPR 2020
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen 等CVPR 2021
相关 Paper
- Gaze Target Detection by Merging Human Attention and Activity CuesYaokun Yang, Yihan Yin, Feng LuAAAI 2024 · 被引用 6 次
- Looking here or there? Gaze Following in 360-Degree ImagesYunhao Li, Wei Shen, Zhongpai Gao, Yucheng Zhu 等ICCV 2021 · 被引用 24 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo 等CVPR 2022 · 被引用 69 次
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu 等CVPR 2023
