GaTector: A Unified Framework for Gaze Object Prediction
Binglu Wang, Tao Hu, Baoshan Li, Xiaojuan Chen, Zhijie Zhang
摘要
Gaze object prediction is a newly proposed task that aims to discover the objects being stared at by humans. It is of great application significance but still lacks a unified solution framework. An intuitive solution is to incorporate an object detection branch into an existing gaze prediction method. However, previous gaze prediction methods usually use two different networks to extract features from scene image and head image, which would lead to heavy network architecture and prevent each branch from joint optimization. In this paper, we build a novel framework named GaTector to tackle the gaze object prediction problem in a unified way. Particularly, a specific-general-specific (SGS) feature extractor is firstly proposed to utilize a shared backbone to extract general features for both scene and head images. To better consider the specificity of inputs and tasks, SGS introduces two input-specific blocks before the shared backbone and three task-specific blocks after the shared backbone. Specifically, a novel Defocus layer is designed to generate object-specific features for the object detection task without losing information or requiring extra computations. Moreover, the energy aggregation loss is introduced to guide the gaze heatmap to concentrate on the stared box. In the end, we propose a novel wUoC metric that can reveal the difference between boxes even when they share no overlapping area. Extensive experiments on the GOO dataset verify the superiority of our method in all three tracks, i.e. object detection, gaze estimation, and gaze object prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Robust Region Feature Synthesizer for Zero-Shot Object DetectionPeiliang Huang, Junwei Han, De Cheng, Dingwen ZhangCVPR 2022 · 被引用 50 次
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 被引用 41 次
- TransGOP: Transformer-Based Gaze Object PredictionBinglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia 等AAAI 2024 · 被引用 8 次
- FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationDaosong Hu, Mingyue Cui, Kai HuangCVPR 2025
- Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersFiona Ryan, Ajay Bati, Sangmin Lee, Daniel Bolya 等CVPR 2025
它引用的顶会 Paper10
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- Detail-Preserving Transformer for Light Field Image Super-resolutionShunzhou Wang, Tianfei Zhou, Yao Lu, Huijun DiAAAI 2022 · 被引用 131 次
- Generalizing Gaze Estimation with Outlier-guided Collaborative AdaptationYunfei Liu, Ruicong Liu, Haofei Wang, Feng LuICCV 2021 · 被引用 80 次
- Colar: Effective and Efficient Online Action Detection by Consulting ExemplarsLe Yang, Junwei Han, Dingwen ZhangCVPR 2022 · 被引用 55 次
相关 Paper
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo 等CVPR 2022 · 被引用 69 次
- Glance and Gaze: Inferring Action-Aware Points for One-Stage Human-Object Interaction DetectionXubin Zhong, Xian Qu, Changxing Ding, Dacheng TaoCVPR 2021
- Cascade-DETR: Delving into High-Quality Universal Object DetectionMingqiao Ye, Lei Ke, Siyuan Li, Yu-Wing Tai 等ICCV 2023 · 被引用 62 次
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard 等NeurIPS 2024 · 被引用 24 次
