Motion-Guided Masking for Spatiotemporal Representation Learning
David Fan, Jue Wang, Shuai Liao, Yi Zhu, Vimal Bhat, Hector J. Santos-Villalobos, Rohith MV, Xinyu Li
Abstract
Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video understanding. This suggests that the random masking strategy that is inherited from the image MAE is less effective for video MAE. This motivates the design of a novel masking algorithm that can more efficiently make use of video saliency. Specifically, we propose a motion-guided masking algorithm (MGM) which leverages motion vectors to guide the position of each mask over time. Crucially, these motion-based correspondences can be directly obtained from information stored in the compressed format of the video, which makes our method efficient and scalable. On two challenging large-scale video benchmarks (Kinetics-400 and Something-Something V2), we equip video MAE with our MGM and achieve up to +1.3% improvement compared to previous state-of-the-art methods. Additionally, our MGM achieves equivalent performance to previous video MAE using up to 66% fewer training epochs. Lastly, we show that MGM generalizes better to downstream transfer learning and domain adaptation tasks on the UCF101, HMDB51, and Diving48 datasets, achieving up to +4.9% improvement compared to baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53f327f9-077a-45c3-b115-e11aae19d4b5Cited by top-tier papers15
- Video Token Merging for Long Video UnderstandingSeon-Ho Lee, Jue Wang, Zhikang Zhang, David Fan et al.NeurIPS 2024 · 21 citations
- EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal TokensSunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju HwangICML 2024 · 16 citations
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei et al.CVPR 2026 · 13 citations
- Multi-view Masked Contrastive Representation Learning for Endoscopic Video AnalysisKai Hu, Ye Xiao, Yuan Zhang, Xieping GaoNeurIPS 2024 · 13 citations
- Causal-JEPA: Learning World Models through Object-Level Latent MaskingHeejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun et al.ICML 2026 · 7 citations
Builds on28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- MGMAE: Motion Guided Masking for Video Masked AutoencodingBingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao et al.ICCV 2023 · 58 citations
- AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked AutoencodersWele Gedara Chaminda Bandara, Naman Patel, Ali Gholami, Mehdi Nikkhah et al.CVPR 2023
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 2 citations
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingLimin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong et al.CVPR 2023
