MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze Prediction
Anshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard, Jean-Marc Odobez
摘要
Gaze following and social gaze prediction are fundamental tasks providing insights into human communication behaviors, intent, and social interactions. Most previous approaches addressed these tasks separately, either by designing highly specialized social gaze models that do not generalize to other social gaze tasks or by considering social gaze inference as an ad-hoc post-processing of the gaze following task. Furthermore, the vast majority of gaze following approaches have proposed models that can handle only one person at a time and are static, therefore failing to take advantage of social interactions and temporal dynamics. In this paper, we address these limitations and introduce a novel framework to jointly predict the gaze target and social gaze label for all people in the scene. It comprises (i) a temporal, transformer-based architecture that, in addition to frame tokens, handles person-specific tokens capturing the gaze information related to each individual; (ii) a new dataset, VSGaze, built from multiple gaze following and social gaze datasets by extending and validating head detections and tracks, and unifying annotation types. We demonstrate that our model can address and benefit from training on all tasks jointly, achieving state-of-the-art results for multi-person gaze following and social gaze prediction. Our annotations and code will be made publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-View Gaze Target EstimationQiaomu Miao, Vivek Raju Golani, Jingyi Xu, Progga Paromita Dutta 等ICCV 2025 · 被引用 4 次
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng 等CVPR 2026 · 被引用 3 次
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai 等NeurIPS 2025 · 被引用 3 次
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction DetectionDongkeun Kim, Minsu Cho, Suha KwakNeurIPS 2025 · 被引用 1 次
- Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following LabelsPierre Vuillecard, Jean-Marc OdobezCVPR 2025
它引用的顶会 Paper15
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
- Understanding Human Gaze Communication by Spatio-Temporal Graph ReasoningLifeng Fan, Wenguan Wang, Song-Chun Zhu, Xinyu Tang 等ICCV 2019 · 被引用 124 次
相关 Paper
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo 等CVPR 2022 · 被引用 69 次
- Sharingan: A Transformer Architecture for Multi-Person Gaze FollowingSamy Tafasca, Anshul Gupta, Jean-Marc OdobezCVPR 2024 · 被引用 15 次
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- GazeOnce: Real-Time Multi-Person Gaze EstimationMingfang Zhang, Yunfei Liu, Feng LuCVPR 2022 · 被引用 30 次
