De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation
Yunfeng Xiao, Xiaowei Bai, Baojun Chen, Hao Su, Hao He, Liang Xie, Erwei Yin
Abstract
3D Gaze estimation is a challenging task due to two main issues. First, existing methods focus on analyzing dense features (e.g., large pixel regions), which are sensitive to local noise (e.g., light spots, blurs) and result in increased computational complexity. Second, an eyeball model can correspond multiple gaze directions, and the entangled representation between gazes and models increases the learning difficulty. To address these issues, we propose De 2 Gaze, a lightweight and accurate model-aware 3D gaze estimation method. In De 2 Gaze, we introduce two key innovations for deformable and decoupled representation learning. Specifically, first, we propose a deformable sparse attention mechanism that can adapt sparse sampling points to attention areas to avoid local noise influences. Second, we propose a spatial decoupling network with a dual-branch decoding architecture to disentangle invariant (e.g., eyeball radius, position) and variable (e.g., gaze, pupil, iris) features from the latent space. Compared to existing methods, De 2 Gaze requires fewer sparse features, and achieves faster convergence speed, lower computational complexity, and higher accuracy in 3D gaze estimation. Qualitative and quantitative experiments demonstrate that De 2 Gaze achieves stateof-the-art accuracy and high-quality semantic segmentation for 3D gaze estimation on the TEyeD dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8e3bc2c-cc4d-429d-9123-5e11f4a0c8f1Cited by top-tier papers2
- 3DPE-Gaze: Unlocking the Potential of 3D Facial Priors for Generalized Gaze EstimationYangshi Ge, Yiwei Bao, Feng LuNeurIPS 2025 · 1 citation
- Seeing the Unseen: Physics-as-Representation for Generalizable Gaze PerceptionYunfeng Xiao, Xiaowei Bai, Hao Su, Hao He et al.ICML 2026
Builds on7
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
Related papers
- Dual Attention Guided Gaze Target Detection in the WildYi Fang, Jiapeng Tang, Wang Shen, Wei Shen et al.CVPR 2021
- Cross-Encoder for Unsupervised Gaze Representation LearningYunjia Sun, Jiabei Zeng, Shiguang Shan, Xilin ChenICCV 2021 · 40 citations
- Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball RotationYoungChan Choi, HengFei Wang, YiHua Cheng, Boeun Kim et al.ACM MM 2025
- Unsupervised Gaze Representation Learning from Multi-view Face ImagesYiwei Bao, Feng LuCVPR 2024
- GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context EncodingYuki Kawana, Shintaro Shiba, Quan Kong, Norimasa KoboriCVPR 2025
