Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark Tracking
Congcong Zhu, Xiaoqiang Li, Jide Li, Guangtai Ding, Weiqin Tong
Abstract
Diversity of training data significantly affects tracking robustness of model under unconstrained environments. However, existing labeled datasets for facial landmark tracking tend to be large but not diverse, and manually annotating the massive clips of new diverse videos is extremely expensive. To address these problems, we propose a Spatial-Temporal Knowledge Integration (STKI) approach. Unlike most existing methods which rely heavily on labeled data, STKI exploits supervisions from unlabeled data. Specifically, STKI integrates spatial-temporal knowledge from massive unlabeled videos, which has several orders of magnitude more than existing labeled video data on the diversity, for robust tracking. Our framework includes a self-supervised tracker and an image-based detector for tracking initialization. To avoid the distortion of facial shape, the tracker leverages adversarial learning to introduce facial structure prior and temporal knowledge into cycle-consistency tracking. Meanwhile, we design a graph-based knowledge distillation method, which distills the knowledge from tracking and detection results, to improve the generalization of the detector. The fine-tuned detector can provide tracker on unconstrained videos with high-quality tracking initialization. Extensive experimental results show that the proposed method achieves state-of-the-art performance on comprehensive evaluation datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3004cfc2-3361-4b3e-bcee-37a865927c55Cited by top-tier papers1
Ask how each one uses itRelated papers
- Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled DataShi Yin, Shangfei Wang, Xiaoping Chen, Enhong ChenACM MM 2020 · 7 citations
- Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosCongcong Zhu, Hao Liu, Zhenhua Yu, Xuehong SunAAAI 2020 · 11 citations
- Unsupervised Learning of Accurate Siamese TrackingQiuhong Shen, Lei Qiao, Jinyang Guo, Peixia Li et al.CVPR 2022 · 73 citations
- PrefAce: Face-Centric Pretraining with Self-Structure Aware DistillationSiyuan Hu, Zheng Wang, Peng Hu, Xi Peng et al.AAAI 2024 · 2 citations
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
