VideoLucy: Deep Memory Backtracking for Long Video Understanding
Jialong Zuo, Yongtai Deng, Lingdong Kong, Jingkang Yang, Rui Jin, Yiwei Zhang, Nong Sang, Liang Pan, Ziwei Liu, Changxin Gao
摘要
Recent studies have shown that agent-based systems leveraging large language models (LLMs) for key information retrieval and integration have emerged as a promising approach for long video understanding. However, these systems face two major challenges. First, they typically perform modeling and reasoning on individual frames, struggling to capture the temporal context of consecutive frames. Second, to reduce the cost of dense frame-level captioning, they adopt sparse frame sampling, which risks discarding crucial information. To overcome these limitations, we propose VideoLucy, a deep memory backtracking framework for long video understanding. Inspired by the human recollection process from coarse to fine, VideoLucy employs a hierarchical memory structure with progressive granularity. This structure explicitly defines the detail level and temporal scope of memory at different hierarchical depths. Through an agent-based iterative backtracking mechanism, VideoLucy systematically mines video-wide, question-relevant deep memories until sufficient information is gathered to provide a confident answer. This design enables effective temporal understanding of consecutive frames while preserving critical details. In addition, we introduce EgoMem, a new benchmark for long video understanding. EgoMem is designed to comprehensively evaluate a model's ability to understand complex events that unfold over time and capture fine-grained details in extremely long videos. Extensive experiments demonstrate the superiority of VideoLucy. Built on open-source models, VideoLucy significantly outperforms state-of-the-art methods on multiple long video understanding benchmarks, achieving performance even surpassing the latest proprietary models such as GPT-4o. Our code and dataset will be made publicly at https://videolucy.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video UnderstandingYufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding 等CVPR 2026 · 被引用 24 次
- Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic SearchXinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong 等CVPR 2026 · 被引用 8 次
- LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life TasksHengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen 等CVPR 2026 · 被引用 4 次
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question AnsweringYuanhao Zou, Shengji Jin, Andong Deng, Youpeng Zhao 等ICLR 2026
- META: Meta Evolution of Tool Trajectory Adaptation for Long-Video UnderstandingJing Huang, Luyuan Chen, Zhijie Xu, Yadong Li 等CVPR 2026
它引用的顶会 Paper33
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng 等ICLR 2026 · 被引用 172 次
相关 Paper
- DrVideo: Document Retrieval Based Long Video UnderstandingZiyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 等CVPR 2025
- VideoSeek: Long-Horizon Video Agent with Tool-Guided SeekingJingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun 等CVPR 2026 · 被引用 15 次
- VideoChat-A1: Thinking with Long Videos by Chain-of-Shot ReasoningZikang Wang, Boyu Chen, Zhengrong Yue, Yi Wang 等AAAI 2026 · 被引用 25 次
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long VideosZiyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon 等CVPR 2025
- VideoBrain: Learning Adaptive Frame Sampling for Long Video UnderstandingJunbo Zou, Ziheng Huang, Shengjie Zhang, Liwen Zhang 等ICML 2026 · 被引用 4 次
