Online Generic Event Boundary Detection
Hyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son, Jonghyun Choi
Abstract
Generic Event Boundary Detection (GEBD) aims to interpret long-form videos through the lens of human perception. However, current GEBD methods require processing complete video frames to make predictions, unlike humans processing data online and in real-time. To bridge this gap, we introduce a new task, Online Generic Event Boundary Detection (On-GEBD), aiming to detect boundaries of generic events immediately in streaming videos. This task faces unique challenges of identifying subtle, taxonomyfree event changes in real-time, without the access to future frames. To tackle these challenges, we propose a novel On-GEBD framework, ESTimator, inspired by Event Segmentation Theory (EST) [50] which explains how humans segment ongoing activity into events by leveraging the discrepancies between predicted and actual information. Our framework consists of two key components: the Consistent Event Anticipator (CEA), and the Online Boundary Discriminator (OBD). Specifically, the CEA generates a prediction of the future frame reflecting current event dynamics based solely on prior frames. Then, the OBD measures the prediction error and adaptively adjusts the threshold using statistical tests on past errors to capture diverse, subtle event transitions. Experimental results demonstrate that ESTimator outperforms all baselines adapted from recent online video understanding models and achieves performance comparable to prior offline-GEBD methods on the Kinetics-GEBD and TAPOS datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 159df2e2-a420-49ef-b817-32d21a9ec1abBuilds on26
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
Related papers
- Generic Event Boundary Detection: A Benchmark for Event SegmentationMike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram et al.ICCV 2021 · 91 citations
- Generic Event Boundary Detection via Denoising DiffusionJaejun Hwang, Dayoung Gong, Manjin Kim, Minsu ChoICCV 2025
- Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary DetectionJiaqi Tang, Zhaoyang Liu, Chen Qian, Wayne Wu et al.CVPR 2022 · 15 citations
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 26 citations
- Rethinking the Architecture Design for Efficient Generic Event Boundary DetectionZiwei Zheng, Zechuan Zhang, Yulin Wang, Shiji Song et al.ACM MM 2024 · 1 citation
