Generic Event Boundary Detection: A Benchmark for Event Segmentation
Mike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram, Matt Feiszli
摘要
This paper presents a novel task together with a new benchmark for detecting generic, taxonomy-free event boundaries that segment a whole video into chunks. Conventional work in temporal video segmentation and action detection focuses on localizing pre-defined action categories and thus does not scale to generic videos. Cognitive Science has known since last century that humans consistently segment videos into meaningful temporal chunks. This segmentation happens naturally, without pre-defined event categories and without being explicitly asked to do so. Here, we repeat these cognitive experiments on mainstream CV datasets; with our novel annotation guideline which addresses the complexities of taxonomy-free event boundary annotation, we introduce the task of Generic Event Boundary Detection (GEBD) and the new benchmark Kinetics-GEBD. We view GEBD as an important stepping stone towards understanding the video as a whole, and believe it has been previously neglected due to a lack of proper task definition and annotations. Through experiment and human study we demonstrate the value of the annotations. Further, we benchmark supervised and un-supervised GEBD approaches on the TAPOS dataset and our Kinetics-GEBD. We release our annotations and baseline codes at CVPR'21 LOVEU Challenge: https://sites.google.com/ view/loveucvpr21 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon 等ICCV 2023 · 被引用 103 次
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 等CVPR 2024 · 被引用 95 次
- Equivariant Similarity for Vision-Language Foundation ModelsTan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin 等ICCV 2023 · 被引用 67 次
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal UnderstandingYunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng 等AAAI 2025 · 被引用 29 次
它引用的顶会 Paper8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 被引用 109 次
相关 Paper
- Generic Event Boundary Detection via Denoising DiffusionJaejun Hwang, Dayoung Gong, Manjin Kim, Minsu ChoICCV 2025
- Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary DetectionJiaqi Tang, Zhaoyang Liu, Chen Qian, Wayne Wu 等CVPR 2022 · 被引用 15 次
- Online Generic Event Boundary DetectionHyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son 等ICCV 2025
- End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionCongcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong 等CVPR 2022 · 被引用 18 次
- Rethinking the Architecture Design for Efficient Generic Event Boundary DetectionZiwei Zheng, Zechuan Zhang, Yulin Wang, Shiji Song 等ACM MM 2024 · 被引用 1 次
