Sharingan: A Transformer Architecture for Multi-Person Gaze Following
Samy Tafasca, Anshul Gupta, Jean-Marc Odobez
Abstract
Gaze is a powerful form of non-verbal communication that humans develop from an early age. As such, modeling this behavior is an important task that can benefit a broad set of application domains ranging from robotics to sociology. In particular, the gaze following task in computer vision is defined as the prediction of the 2D pixel coordinates where a person in the image is looking. Previous attempts in this area have primarily centered on CNN-based architectures, but they have been constrained by the need to process one person at a time, which proves to be highly inefficient. In this paper, we introduce a novel and effective multi-person transformer-based architecture for gaze prediction. While there exist prior works using transformers for multi-person gaze prediction [38], [39], they use a fixed set of learnable embeddings to decode both the person and its gaze target, which requires a matching step afterward to link the predictions with the annotations. Thus, it is difficult to quantitatively evaluate these methods reliably with the available benchmarks, or integrate them into a larger human behavior understanding system. Instead, we are the first to propose a multi-person transformer-based architecture that maintains the original task formulation and ensures control over the people fed as input. Our main contribution lies in encoding the person-specific information into a single controlled token to be processed alongside image tokens and using its output for prediction based on a novel multiscale decoding mechanism. Our new architecture achieves state-of-the-art results on the GazeFollow, VideoAttentionTarget, and ChildPlay datasets and outperforms comparable multi-person architectures with a notable margin. Our code, checkpoints, and data extractions will be made publicly available soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ed5995d-1564-4864-801f-a848b64cfd25Cited by top-tier papers6
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard et al.NeurIPS 2024 · 24 citations
- Toward Semantic Gaze Target DetectionSamy Tafasca, Anshul Gupta, Victor Bros, Jean-Marc OdobezNeurIPS 2024 · 16 citations
- Multi-View Gaze Target EstimationQiaomu Miao, Vivek Raju Golani, Jingyi Xu, Progga Paromita Dutta et al.ICCV 2025 · 4 citations
- Gaze Target Estimation Anywhere with ConceptsXu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou et al.CVPR 2026 · 3 citations
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai et al.NeurIPS 2025 · 3 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo et al.CVPR 2022 · 69 citations
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 41 citations
Related papers
- Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersFiona Ryan, Ajay Bati, Sangmin Lee, Daniel Bolya et al.CVPR 2025
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 38 citations
- Glance-and-Gaze Vision TransformerQihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu et al.NeurIPS 2021 · 91 citations
- TransGOP: Transformer-Based Gaze Object PredictionBinglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia et al.AAAI 2024 · 8 citations
- Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersZhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue et al.CVPR 2024
