EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwang
Abstract
Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memory in predicting uninformative tokens/frames due to random masking strategies. (e.g., over 16 nodes with 128 NVIDIA A100 GPUs (Feichtenhofer et al., 2022)). To resolve this issue, we exploit the unequal information density among the patches in videos and propose EVEREST, a surprisingly efficient MVA approach for video representation learning that finds tokens containing rich motion features and discards uninformative ones during both pre-training and fine-tuning. We further present an information-intensive frame selection strategy that allows the model to focus on informative and causal frames with minimal redundancy. Our method significantly reduces the computation and memory requirements of MVA, enabling the pre-training and fine-tuning on a single machine with 8 GPUs while achieving comparable performance to computation-and memory-heavy baselines on multiple benchmarks and the uncurated Ego4D dataset. We hope that our work contributes to reducing the barrier to further research on video understanding. Codes are available at: https://github.com/sunilhoho/EVEREST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9c1c06c-5f10-42f8-95cb-955f7d2b68bdCited by top-tier papers11
- Don't Look Twice: Faster Video Transformers with Run-Length TokenizationRohan Choudhury, Guanglei Zhu, Sihan Liu, Koichiro Niinuma et al.NeurIPS 2024 · 56 citations
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- Make Your Training Flexible: Towards Deployment-Efficient Video ModelsChenting Wang, Kunchang Li, Tianxiang Jiang, Xiangyu Zeng et al.ICCV 2025 · 8 citations
- Extending Video Masked Autoencoders to 128 framesNitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde et al.NeurIPS 2024 · 6 citations
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 3 citations
Builds on26
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Mask Again: Masked Knowledge Distillation for Masked Video ModelingXiaojie Li, Shaowei He, Jianlong Wu, Yue Yu et al.ACM MM 2023 · 3 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Progressively Compressed Auto-Encoder for Self-supervised Representation LearningJin Li, Yaoming Wang, Xiaopeng Zhang, Yabo Chen et al.ICLR 2023
- Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionQiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu et al.ACM MM 2023 · 9 citations
- MGMAE: Motion Guided Masking for Video Masked AutoencodingBingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao et al.ICCV 2023 · 58 citations
