Sharingan: A Transformer Architecture for Multi-Person Gaze Following
Samy Tafasca, Anshul Gupta, Jean-Marc Odobez
摘要
Gaze is a powerful form of non-verbal communication that humans develop from an early age. As such, modeling this behavior is an important task that can benefit a broad set of application domains ranging from robotics to sociology. In particular, the gaze following task in computer vision is defined as the prediction of the 2D pixel coordinates where a person in the image is looking. Previous attempts in this area have primarily centered on CNN-based architectures, but they have been constrained by the need to process one person at a time, which proves to be highly inefficient. In this paper, we introduce a novel and effective multi-person transformer-based architecture for gaze prediction. While there exist prior works using transformers for multi-person gaze prediction [38], [39], they use a fixed set of learnable embeddings to decode both the person and its gaze target, which requires a matching step afterward to link the predictions with the annotations. Thus, it is difficult to quantitatively evaluate these methods reliably with the available benchmarks, or integrate them into a larger human behavior understanding system. Instead, we are the first to propose a multi-person transformer-based architecture that maintains the original task formulation and ensures control over the people fed as input. Our main contribution lies in encoding the person-specific information into a single controlled token to be processed alongside image tokens and using its output for prediction based on a novel multiscale decoding mechanism. Our new architecture achieves state-of-the-art results on the GazeFollow, VideoAttentionTarget, and ChildPlay datasets and outperforms comparable multi-person architectures with a notable margin. Our code, checkpoints, and data extractions will be made publicly available soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard 等NeurIPS 2024 · 被引用 24 次
- Toward Semantic Gaze Target DetectionSamy Tafasca, Anshul Gupta, Victor Bros, Jean-Marc OdobezNeurIPS 2024 · 被引用 16 次
- Multi-View Gaze Target EstimationQiaomu Miao, Vivek Raju Golani, Jingyi Xu, Progga Paromita Dutta 等ICCV 2025 · 被引用 4 次
- Gaze Target Estimation Anywhere with ConceptsXu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou 等CVPR 2026 · 被引用 3 次
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo 等CVPR 2022 · 被引用 69 次
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 被引用 41 次
相关 Paper
- Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersFiona Ryan, Ajay Bati, Sangmin Lee, Daniel Bolya 等CVPR 2025
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- Glance-and-Gaze Vision TransformerQihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu 等NeurIPS 2021 · 被引用 91 次
- TransGOP: Transformer-Based Gaze Object PredictionBinglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia 等AAAI 2024 · 被引用 8 次
- Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersZhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue 等CVPR 2024
