Boundary-sensitive Pre-training for Temporal Localization in Videos
Mengmeng Xu, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martínez, Xiatian Zhu, Li Zhang, Bernard Ghanem, Tao Xiang
Abstract
Many video analysis tasks require temporal localization for the detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is due to large scale annotation of temporal boundaries in untrimmed videos being expensive. Therefore, no suitable datasets exist that enable pre-training in a manner sensitive to temporal boundaries. In this paper for the first time, we investigate model pre-training for temporal localization by introducing a novel boundary-sensitive pretext (BSP) task. Instead of relying on costly manual annotations of temporal boundaries, we propose to synthesize temporal boundaries in existing video action classification datasets. By defining different ways of synthesizing boundaries, BSP can then be simply conducted in a self-supervised manner via the classification of the boundary types. This enables the learning of video representations that are much more transferable to downstream temporal localization tasks. Extensive experiments show that the proposed BSP is superior and complementary to the existing action classification-based pre-training counterpart, and achieves new state-of-the-art performance on several temporal localization tasks. Please visit our website for more details https://frostinassiky.github.io/bsp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- Video Self-Stitching Graph Network for Temporal Action LocalizationChen Zhao, Ali K. Thabet, Bernard GhanemICCV 2021 · 179 citations
- Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationFa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan et al.ACM MM 2021 · 104 citations
- An Empirical Study of End-to-End Temporal Action DetectionXiaolong Liu, Song Bai, Xiang BaiCVPR 2022 · 72 citations
- Unsupervised Pre-training for Temporal Action Localization TasksCan Zhang, Tianyu Yang, Junwu Weng, Meng Cao et al.CVPR 2022 · 56 citations
Builds on20
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
Related papers
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Test-Time Zero-Shot Temporal Action LocalizationBenedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang et al.CVPR 2024
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Weakly-Guided Self-Supervised Pretraining for Temporal Activity DetectionKumara Kahatapitiya, Zhou Ren, Haoxiang Li, Zhenyu Wu et al.AAAI 2023 · 7 citations
- PreDet: Large-scale weakly supervised pre-training for detectionVignesh Ramanathan, Rui Wang, Dhruv MahajanICCV 2021 · 14 citations
