Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled Data
Shi Yin, Shangfei Wang, Xiaoping Chen, Enhong Chen
Abstract
Current work of facial landmark tracking usually requires large amounts of fully annotated facial videos to train a landmark tracker. To relieve the burden of manual annotations, we propose a novel facial landmark tracking method that makes full use of unlabeled facial videos by exploiting both self-supervised and semi-supervised learning mechanisms. First, self-supervised learning is adopted for representation learning from unlabeled facial videos. Specifically, a facial video and its shuffled version are fed into a feature encoder and a classifier. The feature encoder is used to learn visual representations, and the classifier distinguishes the input videos as the original or the shuffled ones. The feature encoder and the classifier are trained jointly. Through self-supervised learning, the spatial and temporal patterns of a facial video are captured at representation level. After that, the facial landmark tracker, consisting of the pre-trained feature encoder and a regressor, is trained semi-supervisedly. The consistencies among the tracking results of the original, the inverse and the disturbed facial sequences are exploited as the constraints on the unlabeled facial videos, and the supervised loss is adopted for the labeled videos. Through semi-supervised end-to-end training, the tracker captures sequential patterns inherent in facial videos despite small amount of manual annotations. Experiments on two benchmark datasets show that the proposed framework outperforms state-of-the-art semi-supervised facial landmark tracking methods, and also achieves advanced performance compared to fully supervised facial landmark tracking methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 29d232a5-66b4-41c7-9db8-141845d059ffRelated papers
- Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark TrackingCongcong Zhu, Xiaoqiang Li, Jide Li, Guangtai Ding et al.ACM MM 2020 · 10 citations
- Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosCongcong Zhu, Hao Liu, Zhenhua Yu, Xuehong SunAAAI 2020 · 11 citations
- Exploiting Invariance of Mining Facial LandmarksJiangming Shi, Zixian Gao, Hao Liu, Zekuan Yu et al.ACM MM 2021 · 2 citations
- LAFS: Landmark-Based Facial Self-Supervised Learning for Face RecognitionZhonglin Sun, Chen Feng, Ioannis Patras, Georgios TzimiropoulosCVPR 2024 · 17 citations
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
