Mask Again: Masked Knowledge Distillation for Masked Video Modeling
Xiaojie Li, Shaowei He, Jianlong Wu, Yue Yu, Liqiang Nie, Min Zhang
摘要
Masked video modeling has shown remarkable performance in downstream tasks by predicting masked video tokens from visible ones. However, training models from scratch on large-scale unlabeled data remains computationally challenging and time-consuming. Moreover, the commonly used random-based sampling techniques may lead to the selection of redundant or low-information regions, hindering the model from learning discriminative representations within the limited training epochs. To achieve efficient pre-training, we propose MaskAgain, an efficient feature-based knowledge distillation framework for masked video pre-training that facilitates knowledge transfer from a pre-trained teacher model to a student model. In contrast to previous approaches that align all visible token features with the teacher model at output layers, MaskAgain adopts a selective approach by masking visible tokens again at both the hidden and output layers of the transformer block. Attention mechanisms are utilized for informative feature selection. At the hidden level, attention maps generated by the transformer's multi-head attention structure are utilized to select crucial token information at both temporally-global and temporally-local levels. Additionally, at the output level, an activation-based attention map is generated using token features, enabling us to focus on important tokens while preserving feature similarity and the relationship matrix similarity between patches. Extensive experimental results show that MaskAgain achieves comparable or even better performance than existing methods on benchmark datasets with much fewer training epochs and much less memory, which demonstrates that MaskAgain allows for efficient pre-training of accurate video models, reducing computational resources and training time significantly. Code is released at https://github.com/xiaojieli0903/MaskAgain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等NeurIPS 2024 · 被引用 104 次
- Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic ManipulationQi Lv, Hao Li, Xiang Deng, Rui Shao 等CVPR 2025
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 等CVPR 2025
- DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any ArchitectureQianlong Xiang, Miao Zhang, Yuzhang Shang, Jianlong Wu 等CVPR 2025
- CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuningYibo Yang, Xiaojie Li, Zhongzhu Zhou, Shuaiwen Song 等NeurIPS 2024
它引用的顶会 Paper40
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
相关 Paper
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen 等CVPR 2023
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson 等CVPR 2026 · 被引用 9 次
- Asymmetric Masked Distillation for Pre-Training Small Foundation ModelsZhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu 等CVPR 2024 · 被引用 6 次
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang 等ICCV 2023 · 被引用 266 次
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingLimin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong 等CVPR 2023
