ReWind: Understanding Long Videos with Instructed Learnable Memory
Anxhelo Diko, Tinghuai Wang, Wassim Swaileh, Shiyan Sun, Ioannis Patras
摘要
Vision-Language Models (VLMs) are crucial for applications requiring integrated understanding textual and visual information. However, existing VLMs struggle with long videos due to computational inefficiency, memory limitations, and difficulties in maintaining coherent understanding across extended sequences. To address these challenges, we introduce ReWind, a novel memory-based VLM designed for efficient long video understanding while preserving temporal fidelity. ReWind operates in a two-stage framework. In the first stage, ReWind maintains a dynamic learnable memory module with a novel read-perceive-write cycle that stores and updates instruction-relevant visual information as the video unfolds. This module utilizes learnable queries and cross-attentions between memory contents and the input stream, ensuring low memory requirements by scaling linearly with the number of tokens. In the second stage, we propose an adaptive frame selection mechanism guided by the memory content to identify instruction-relevant key moments. It enriches the memory representations with detailed spatial information by selecting a few high-resolution frames, which are then combined with the memory contents and fed into a Large Language Model (LLM) to generate the final answer. We empirically demonstrate ReWind's superior performance in visual question answering (VQA) and temporal grounding tasks, surpassing previous methods on long video benchmarks. Notably, ReWind achieves a +13% score gain and a +12% accuracy improvement on the MovieChat-1K VQA dataset and an +8% mIoU increase on Charades-STA for temporal grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic SearchXinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong 等CVPR 2026 · 被引用 8 次
- Agentic Very Long Video UnderstandingAniket Rege, Arka Sadhu, Yuliang Li, Kejie Li 等ACL 2026 · 被引用 8 次
- Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic MemoryRongjie Jiang, Jianwei Wang, Gengda Zhao, Chengyang Luo 等KDD 2026 · 被引用 5 次
- Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video UnderstandingHoulun Chen, Xin Wang, Guangyao Li, Yuwei Zhou 等SIGIR 2026
它引用的顶会 Paper12
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized AttentionRenrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou 等ICLR 2024 · 被引用 174 次
相关 Paper
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingBo He, Hengduo Li, Young Kyun Jang, Menglin Jia 等CVPR 2024
- AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory ReductionYuanbin Man, Ying Huang, Chengming Zhang, Bingzhe Li 等CVPR 2025
- Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video ProcessingYudong Liu, Jingwei Sun, Yueqian Lin, Jianyi Zhang 等ICCV 2025 · 被引用 22 次
