WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMs
Yulin Zhang, Cheng Shi, Sibei Yang
摘要
Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them ill-suited for streaming settings where frames arrive sequentially and future observations are inaccessible. We diagnose a core limitation of current Video-LLMs, namely Time-Agnosticism, in which videos are treated as an unordered bag of evidence rather than a causally ordered sequence, yielding two failures in streams: temporal order ambiguity, in which the model cannot follow or reason over the correct chronological order, and past-current focus blindness where it fails to distinguish present observations from accumulated history. We present WeaveTime, a simple, efficient, and model-agnostic framework that first teaches order and then uses order. We introduce a lightweight Temporal Reconstruction objective, namely the Streaming Order Perception enhancement, which instills order-aware representations with minimal finetuning and no specialized streaming data. At inference, a Past-Current Dynamic Focus Cache performs uncertainty-triggered, coarse-to-fine retrieval, expanding history only when needed. Plugged into exsiting Video-LLM without architectural changes, WeaveTime delivers consistent gains on representative streaming benchmarks, improving accuracy while reducing latency. These results establish WeaveTime as a practical path toward time-aware stream Video-LLMs under strict online, time-causal constraints. Code and weights are publicly available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li 等NeurIPS 2025 · 被引用 79 次
相关 Paper
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao 等NeurIPS 2025 · 被引用 5 次
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?Junbo Niu, Yifei Li, Ziyang Miao, Chunjiang Ge 等CVPR 2025
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video UnderstandingHaowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng 等ACL 2026 · 被引用 17 次
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language ModelsJialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu 等CVPR 2026 · 被引用 6 次
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 等AAAI 2026 · 被引用 1 次
