Progressive Mask Distillation for Self-supervised Video Representation
Kewei Wu, Chong Liang, Zhao Xie, Dan Guo
Abstract
Masked visual modeling is a self-supervised learning task that does not use visual annotations. It aims to learn discriminative representations via a mask-reconstruction task. A single mask ratio in reconstruction may fail to capture complex semantics, which motivates dynamic masking strategies. In this work, we propose Progressive Mask Distillation (PMD), which utilizes dynamic mask ratios to facilitate progressive semantic learning from easy to hard. PMD integrates three key components: a progressive student distiller, a difficulty-aware region enhancer, and a cross-layer feature aligner. First, to capture dynamic visual semantics, we design a progressive student distiller that trains multiple student models with progressively increasing mask ratios. The early-phase student (with a low mask ratio) learns easy, low-level semantics from more visible tokens. This learned knowledge then guides the next-phase student (with a higher mask ratio) to capture hard, high-level semantics from fewer visible tokens. This progressive distillation mechanism enhances detail reconstruction at a high mask ratio. Second, to alleviate insufficient learning of semantic regions, we design a difficulty-aware region enhancer. It uses the region reconstruction loss to learn their weights, prioritizing accurate learning of regions with large reconstruction losses. Third, to further bridge the semantic gap across network layers, we design cross-layer feature alignment. This module aligns features across shallow, middle, and deep encoder layers, ensuring that shallow-layer features incorporate semantic information from deeper layers. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the SSv2, K400, UCF-101, and HMDB-51 datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on29
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- BEVT: BERT Pretraining of Video TransformersRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2022 · 200 citations
- Diffusion Models as Masked AutoencodersChen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li et al.ICCV 2023 · 82 citations
Related papers
- Learn More for Food Recognition via Progressive Self-DistillationYaohui Zhu, Linhu Liu, Jiang TianAAAI 2023 · 10 citations
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
- DDAE: Towards Deep Dynamic Vision BERT PretrainingHonghao Chen, Xiangwen Kong, Xiangyu Zhang, Xin Zhao et al.AAAI 2024 · 1 citation
- Exploring Target Representations for Masked AutoencodersXingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin et al.ICLR 2024 · 59 citations
- Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD DatasetsMuhammad Abdullah Jamal, Omid MohareriCVPR 2025
