OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
Zhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang, Haonan Lu, Guanbin Li
摘要
Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis -- small, critical, and easily lost in a desert of redundancy. Enlarging memory only widens the desert; aggressive compression dries up the oasis. The real difficulty lies in discovering where to look, not how much to remember. We therefore introduce OASIS, a novel framework for streaming video reasoning that tackles this challenge through structured, on-demand retrieval. It organizes streaming history into hierarchical events and performs reasoning as controlled refinement -- short-context inference first, followed by semantically grounded retrieval only when uncertainty arises. As the retrieval is driven by high-level intent rather than embedding similarity, the retrieve memory is substantially more accurate and less noisy. Additionally, the mechanism is plug-and-play, training-free, and compatible with any streaming MLLM. Experiments across multiple benchmarks show that OASIS achieves strong gains in long-horizon accuracy and compositional reasoning with far less memory budget.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin 等NeurIPS 2025 · 被引用 164 次
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 等CVPR 2024 · 被引用 95 次
相关 Paper
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video UnderstandingYufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding 等CVPR 2026 · 被引用 24 次
- Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent DecisionsKecheng Zhang, Zongxin Yang, Mingfei Han, Haihong Hao 等ICLR 2026 · 被引用 6 次
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video UnderstandingMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangNeurIPS 2025 · 被引用 62 次
- CogStream: Context-guided Streaming Video Question AnsweringZicheng Zhao, Kangyu Wang, Shijie Li, Rui Qian 等AAAI 2026 · 被引用 3 次
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long VideosZiyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon 等CVPR 2025
