GazeShift: Unsupervised Gaze Estimation and Dataset for VR
Gil Shapira, Ishay Goldin, Evgeny Artyomov, Donghoon Kim, Yosi Keller, Niv Zehngut
Abstract
Gaze estimation is instrumental in modern virtual reality (VR) systems. Despite significant progress in remote-camera gaze estimation, VR gaze research remains constrained by data scarcity, particularly the lack of large-scale, accurately labeled datasets captured with the off-axis camera configurations typical of modern headsets. Gaze annotation is difficult since fixation on intended targets cannot be guaranteed. To address these challenges, we introduce VRGaze, the first large-scale off-axis gaze estimation dataset for VR, comprising 2.1 million near-eye infrared images collected from 68 participants. We further propose GazeShift, an attention-guided unsupervised framework for learning gaze representations without labeled data. Unlike prior redirection-based methods that rely on multi-view or 3D geometry, GazeShift is tailored to near-eye imagery, achieving effective gaze-appearance disentanglement in a compact, real-time model. GazeShift embeddings can be optionally adapted to individual users via lightweight few-shot calibration, achieving a 1.84 mean error on VRGaze. On the remote-camera MPIIGaze dataset, the model achieves a 7.15 person-agnostic error, doing so with 10x fewer parameters and 35x fewer FLOPs than baseline methods. Deployed natively on a VR headset GPU, inference takes only 5 ms. Combined with demonstrated robustness to illumination changes, these results highlight GazeShift as a label-efficient, real-time solution for VR gaze tracking. Project code and the VRGaze dataset are released at https://github.com/gazeshift3/gazeshift
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d543097-56e6-4bdf-8f04-ea0e354e2935Builds on9
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Radi-Eye: Hands-Free Radial Interfaces for 3D Interaction using Gaze-Activated Head-CrossingLudwig Sidenmark, Dominic Potts, Bill Bapisch, Hans GellersenCHI 2021 · 49 citations
- A View on the Viewer: Gaze-Adaptive Captions for VideosKuno Kurzhals, Fabian Göbel, Katrin Angerbauer, Michael Sedlmair et al.CHI 2020 · 42 citations
- Cross-Encoder for Unsupervised Gaze Representation LearningYunjia Sun, Jiabei Zeng, Shiguang Shan, Xilin ChenICCV 2021 · 40 citations
Related papers
- Unsupervised Representation Learning for Gaze EstimationYu Yu, Jean-Marc OdobezCVPR 2020
- Unsupervised Gaze Representation Learning from Multi-view Face ImagesYiwei Bao, Feng LuCVPR 2024
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Hybrid-Domain Adaptative Representation Learning for Gaze EstimationQida Tan, Hongyu Yang, Wenchao DuAAAI 2026
- GazeOnce: Real-Time Multi-Person Gaze EstimationMingfang Zhang, Yunfei Liu, Feng LuCVPR 2022 · 30 citations
