Video Language Model Pretraining with Spatio-temporal Masking
Yue Wu, Zhaobo Qi, Junshu Sun, Yaowei Wang, Qingming Huang, Shuhui Wang
摘要
The development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image features yields superior downstream performance compared to video feature reconstruction. We hypothesize that this performance gap stems from the way how masking strategies influence the model's attention to temporal dynamics. To validate this hypothesis, we performed two sets of experiments that demonstrate that alignment between the masked target and the reconstruction target is crucial for self-supervised video-language learning. Based on these findings, we propose a spatio-temporal masking strategy (STM) for video-language model pretraining that operates across adjacent frames, and a decoder leverages semantic information to enhance the spatiotemporal representations of masked tokens. Thanks to the combination of masking strategy and reconstruction decoder, STM enforces the model to learn spatio-temporal feature representation comprehensively. Experiments in three video understanding downstream tasks validate the superiority of our method. Codes are available here. * Corresponding author. (a) masked targets w./wo. consistent reconstruction targets. (b) Reconstructing image/video feature with different masking strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Relieving the Over-Aggregating Effect in Graph TransformersJunshu Sun, Wanxing Chang, Chenxue Yang, Qingming Huang 等NeurIPS 2025 · 被引用 3 次
- Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language PretrainingWeijun Zhuang, Yuqing Huang, Weikang Meng, Xin Li 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper37
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu 等CVPR 2022 · 被引用 524 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
相关 Paper
- SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningFida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard GhanemCVPR 2025
- MGMAE: Motion Guided Masking for Video Masked AutoencodingBingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao 等ICCV 2023 · 被引用 58 次
- Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationYujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu 等AAAI 2022 · 被引用 20 次
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen 等CVPR 2023
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 被引用 3 次
