Differential Contrastive Training for Gaze Estimation
Lin Zhang, Yi Tian, Xiyun Wang, Wanru Xu, Yi Jin, Yaping Huang
摘要
The complex application scenarios have raised critical requirements for precise and generalizable gaze estimation methods. Recently, the pre-trained CLIP has achieved remarkable performance on various vision tasks, but its potentials have not been fully exploited in gaze estimation. In this paper, we propose a novel Differential Contrastive Training strategy, which boosts gaze estimation performance with the help of the CLIP. Accordingly, a Differential Contrastive Gaze Estimation network (DCGaze) composed of a Visual Appearance-aware branch and a Semantic Differentialaware branch is introduced. The Visual Appearance-aware branch is essentially a primary gaze estimation network and it incorporates an Adaptive Feature-refinement Unit (AFU) and a Doublehead Gaze Regressor (DGR), which both help the primary network to extract informative and gaze-related appearance features. Moreover, the Semantic Difference-aware branch is designed on the basis of the CLIP's text encoder to reveal the semantic difference of gazes. This branch could further empower the Visual Appearance-aware branch with the capability of characterizing the gaze-related semantic information. Extensive experimental results on four challenging datasets over within and cross-domain tasks demonstrate the effectiveness of our DCGaze. The code is available at https:// github.com/ LinZhang-bjtu/ DCGaze.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze EstimationYihua Cheng, Shiyao Huang, Fei Wang, Chen Qian 等AAAI 2020 · 被引用 204 次
- PureGaze: Purifying Gaze Feature for Generalizable Gaze EstimationYihua Cheng, Yiwei Bao, Feng LuAAAI 2022 · 被引用 121 次
相关 Paper
- CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic ModelPengwei Yin, Guanzhong Zeng, Jingjing Wang, Di XieAAAI 2024 · 被引用 29 次
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei 等NeurIPS 2024 · 被引用 70 次
- CLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware PromptingShaoxiang Guo, Qing Cai, Lin Qi, Junyu DongACM MM 2023 · 被引用 10 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang 等CVPR 2023
