Hierarchical Self-supervised Representation Learning for Movie Understanding
Fanyi Xiao, Kaustav Kundu, Joseph Tighe, Davide Modolo
摘要
Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understanding and propose a novel hierarchical self-supervised pretraining strategy that separately pretrains each level of our hierarchical movie understanding model (based on [37]). Specifically, we propose to pretrain the low-level video backbone using a contrastive learning objective, while pretrain the higher-level video contextualizer using an event mask prediction task, which enables the usage of different data sources for pretraining different levels of the hierarchy. We first show that our self-supervised pre-training strategies are effective and lead to improved performance on all tasks and metrics on VidSitu benchmark [37] (e.g., improving on semantic role prediction from 47% to 61% CIDEr scores). We further demonstrate the effectiveness of our contextualized event features on LVU tasks [54], both when used alone and when combined with instance features, showing their complementarity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu 等ICCV 2023 · 被引用 31 次
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 被引用 19 次
- Long-range Multimodal Pretraining for Movie UnderstandingDawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon 等ICCV 2023 · 被引用 15 次
- Large Content And Behavior Models To Understand, Simulate, And Optimize Content And BehaviorAshmit Khandelwal, Aditya Agrawal, Aanisha Bhattacharyya, Yaman Kumar 等ICLR 2024 · 被引用 11 次
- A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero ShotAanisha Bhattacharyya, Yaman Singla, Balaji Krishnamurthy, Rajiv Ratn Shah 等EMNLP 2023 · 被引用 9 次
它引用的顶会 Paper13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 被引用 405 次
相关 Paper
- UniViT: Unifying Image and Video Understanding in One Vision EncoderFeilong Tang, Xiang An, Haolin Yang, Yin Xie 等NeurIPS 2025 · 被引用 3 次
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 被引用 14 次
- SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role LabelingJu-Hee Lee, Je-Won KangCVPR 2024
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu 等ICCV 2023 · 被引用 6 次
- Unsupervised Open-Vocabulary Object Localization in VideosKe Fan, Zechen Bai, Tianjun Xiao, Dominik Zietlow 等ICCV 2023 · 被引用 14 次
