End-to-End Human-Gaze-Target Detection with Transformers
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen
摘要
In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Contrast Everything: A Hierarchical Contrastive Framework for Medical Time-SeriesYihe Wang, Yu Han, Haishuai Wang, Xiang ZhangNeurIPS 2023 · 被引用 106 次
- Saliency in Augmented RealityHuiyu Duan, Wei Shen, Xiongkuo Min, Danyang Tu 等ACM MM 2022 · 被引用 42 次
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 被引用 41 次
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard 等NeurIPS 2024 · 被引用 24 次
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Tracking Pedestrian Heads in Dense CrowdRamana Sundararaman, Cedric De Almeida Braga, Éric Marchand, Julien PettréCVPR 2021
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
- Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With TransformersSixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu 等CVPR 2021
相关 Paper
- Sharingan: A Transformer Architecture for Multi-Person Gaze FollowingSamy Tafasca, Anshul Gupta, Jean-Marc OdobezCVPR 2024 · 被引用 15 次
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen 等CVPR 2021
- Gaze Target Detection by Merging Human Attention and Activity CuesYaokun Yang, Yihan Yin, Feng LuAAAI 2024 · 被引用 6 次
- TransGOP: Transformer-Based Gaze Object PredictionBinglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia 等AAAI 2024 · 被引用 8 次
- HOTR: End-to-End Human-Object Interaction Detection With TransformersBumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim 等CVPR 2021
