Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal Grounding
Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu
摘要
In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike regular video temporal grounding, On-VTG requires the model to make predictions without observing future frames. As online videos are streaming inputs and can go on indefinitely, it is impractical and inefficient to store all historical inputs. The existing OnVTG models employ memory to store recent historical video frame features and predict scores indicating whether the current frame corresponds to the start or end time of the target event. However, these methods lack effective event modeling and cannot retain long-term historical information, leading to low performance. To tackle these challenges, we propose a hierarchical event memory for OnVTG. We propose an event-based OnVTG framework that makes predictions based on event proposals that model event-level information with various durations. To preserve historically valuable event information, we introduce a hierarchical event memory that retains historical events, allowing the model to access both recent and long-term information. To enable the real-time prediction, we further propose a future prediction branch that predicts whether the target event will occur shortly and further regresses the start time of the event. We achieve state-of-the-art performance on the TACoS, Activ-ityNet Captions, and MAD datasets. Code is available at https://github.com/minghangz/OnVTG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OASIS: On-Demand Hierarchical Event Memory for Streaming Video ReasoningZhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang 等CVPR 2026 · 被引用 16 次
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng 等CVPR 2026 · 被引用 4 次
- Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video UnderstandingKe Ma, Jiaqi Tang, Bin Guo, Xueting Han 等ACL 2026
- Temporal-Aware Reasoning Optimization for Video Temporal GroundingMinghang Zheng, Zihao Yin, YI YANG, Yuxin Peng 等ICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis 等ICCV 2019 · 被引用 201 次
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li 等NeurIPS 2021 · 被引用 196 次
相关 Paper
- Temporal Sentence Grounding in Streaming VideosTian Gan, Xiao Wang, Yan Sun, Jianlong Wu 等ACM MM 2023 · 被引用 5 次
- OVG-HQ: Online Video Grounding with Hybrid-Modal QueriesRunhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan 等ICCV 2025 · 被引用 4 次
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- TRACE: Temporal Grounding Video LLM via Causal Event ModelingYongxin Guo, Jingyu Liu, Mingda Li, Qingbin Liu 等ICLR 2025
