Memory Consolidation Enables Long-Context Video Understanding
Ivana Balazevic, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni, Skanda Koppula, Olivier J. Hénaff
Abstract
Most transformer-based video encoders are limited to short temporal contexts due to their quadratic complexity. While various attempts have been made to extend this context, this has often come at the cost of both conceptual and computational complexity. We propose to instead re-purpose existing pre-trained video transformers by simply fine-tuning them to attend to memories derived non-parametrically from past activations. By leveraging redundancy reduction, our memory-consolidated vision transformer (MC-ViT) effortlessly extends its context far into the past and exhibits excellent scaling behavior when learning from longer videos. In doing so, MC-ViT sets a new state-of-the-art in long-context video understanding on EgoSchema, Perception Test, and Diving48, outperforming methods that benefit from orders of magnitude more parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- Temporal Chain of Thought: Long-Video Understanding by Thinking in FramesAnurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi et al.NeurIPS 2025 · 31 citations
- StreamReady: Learning What to Answer and When in Long Streaming VideosShehreen Azad, Vibhav Vineet, Yogesh S. RawatCVPR 2026 · 19 citations
- TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-AlignmentWei Li, Hehe Fan, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2024 · 18 citations
- Flash-Vstream: Efficient Real-Time Understanding for Long Video StreamsHaoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu et al.ICCV 2025 · 15 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan et al.CVPR 2022 · 158 citations
- A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 FramesPinelopi Papalampidi, Skanda Koppula, Shreya Pathak, Justin Chiu et al.CVPR 2024 · 15 citations
- Long-term Leap Attention, Short-term Periodic Shift for Video ClassificationHao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah NgoACM MM 2022 · 12 citations
- Multiview Transformers for Video RecognitionShen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu et al.CVPR 2022 · 279 citations
- One Trajectory, One Token: Grounded Video Tokenization Via Panoptic Sub-Object TrajectoryChenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao et al.ICCV 2025 · 7 citations
