Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labels
Pierre Vuillecard, Jean-Marc Odobez
Abstract
Image GT Supervised (Gaze360) image inference ST-WSGE (Gaze360+GF) video inference ST-WSGE (Gaze360+GF) image inference Figure 1. Significance of ST-WSGE. Our self-training based weakly-supervised framework for robust 3D gaze estimation in real-world conditions (e.g., varying appearance, extreme poses, resolution, and occlusion). All predictions used our image and video agnostic Gaze Transformer (GaT) model. Top row: importance of the training diversity using ST-WSGE and GazeFollow (GF) for generalization compared to standard supervised methods. Bottom row: influence of temporal context between image and video inference. Circles in images represent unit disks where 3D gaze vectors are projected onto the image plane (x, y in yellow) and a top-down view (x, z in blue). Images from VideoAttentionTarget, GFIE, and MPIIFaceGaze datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e259548-3f7d-478c-a633-e5bd13f29383Cited by top-tier papers2
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the WildHongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao et al.NeurIPS 2025 · 15 citations
- Semi-Supervised Gaze Estimation via Disentangled Subspace Contrastive LearningQida Tan, Hongyu Yang, Wenchao DuICML 2026
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
Related papers
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- Weakly-Supervised Physically Unconstrained Gaze EstimationRakshit Sunil Kothari, Shalini De Mello, Umar Iqbal, Wonmin Byeon et al.CVPR 2021
- From Feature to Gaze: A Generalizable Replacement of Linear Layer for Gaze EstimationYiwei Bao, Feng LuCVPR 2024
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo et al.CVPR 2022 · 69 citations
- GazeOnce360: Fisheye-Based 360° Multi-Person Gaze Estimation with Global–Local Feature FusionZhuojiang Cai, Zhenghui Sun, Feng LuCVPR 2026
