MoEGaze: A Mixture of Experts Approach for Generalizable Gaze Estimation
Zheng Liu, Feng Lu
Abstract
Existing gaze estimation models often struggle to generalize to unseen users, primarily due to significant variations in individual appearance. Empirical observations reveal that performance improves when the visual appearance of test subjects closely resembles that of training subjects. Motivated by this, we propose a generalizable gaze estimation framework MoEGaze based on the Mixture of Experts (MoE) architecture. During training, the model extracts appearance features from facial images and uses them to route samples to specialized gaze expert networks, each tailored to a specific subset of appearances. Rather than directly predicting gaze, each expert outputs intermediate gaze features, which are dynamically aggregated according to the input appearance and then mapped to gaze prediction. This dynamic routing design enables the model to effectively adapt to users with diverse appearances, while also facilitating easier training on sub-datasets with smaller appearance variations. Extensive experiments demonstrate that our method achieves superior cross-domain performance compared to existing approaches, with an average improvement of 27.6% across four cross-domain metrics over the baseline. Furthermore, MoEGaze surpasses baselines trained on the full dataset while requiring only 10% of the training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b6948e3-b6ae-4d73-9863-e1fb538b1628Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze EstimationYihua Cheng, Shiyao Huang, Fei Wang, Chen Qian et al.AAAI 2020 · 204 citations
Related papers
- PureGaze: Purifying Gaze Feature for Generalizable Gaze EstimationYihua Cheng, Yiwei Bao, Feng LuAAAI 2022 · 121 citations
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li et al.NeurIPS 2025 · 8 citations
- Generalizing Gaze Estimation with Outlier-guided Collaborative AdaptationYunfei Liu, Ruicong Liu, Haofei Wang, Feng LuICCV 2021 · 80 citations
- CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic ModelPengwei Yin, Guanzhong Zeng, Jingjing Wang, Di XieAAAI 2024 · 29 citations
- Learning a Generalized Gaze Estimator from Gaze-Consistent FeatureMingjie Xu, Haofei Wang, Feng LuAAAI 2023 · 38 citations
