∞-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation
Saul José Rodrigues dos Santos, António Farinhas, Daniel C. McNamee, André F. T. Martins
摘要
Current video-language models struggle with long-video understanding due to limited context lengths and reliance on sparse frame subsampling, often leading to information loss. This paper introduces ∞-VIDEO, which can process arbitrarily long videos through a continuous-time long-term memory (LTM) consolidation mechanism. Our framework augments video Q-formers by allowing them to process unbounded video contexts efficiently and without requiring additional training. Through continuous attention, our approach dynamically allocates higher granularity to the most relevant video segments, forming "sticky" memories that evolve over time. Experiments with Video-LLaMA and VideoChat2 demonstrate improved performance in video question-answering tasks, showcasing the potential of continuoustime LTM mechanisms to enable scalable and training-free comprehension of long videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 被引用 53 次
- FluxMem: Adaptive Hierarchical Memory for Streaming Video UnderstandingYiweng Xie, Bo He, Junke Wang, Xiangyu Zheng 等CVPR 2026 · 被引用 25 次
- A Multi-Agent Perception-Action Alliance for Efficient Long Video ReasoningYichang Xu, Gaowen Liu, Ramana Rao Kompella, Tiansheng Huang 等CVPR 2026 · 被引用 2 次
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen 等ICML 2026
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
相关 Paper
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingBo He, Hengduo Li, Young Kyun Jang, Menglin Jia 等CVPR 2024
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- ∞-former: Infinite Memory TransformerPedro Henrique Martins, Zita Marinho, André F. T. MartinsACL 2022 · 被引用 12 次
- HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video UnderstandingShehreen Azad, Vibhav Vineet, Yogesh Singh RawatCVPR 2025
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory MechanismTao Chen, Kun Zhang, Qiong Wu, Xiao Chen 等CVPR 2026 · 被引用 8 次
