Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos
Sixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo, Yicheng Qian, Shenghua Gao
摘要
Sequential video understanding, as an emerging video understanding task, has driven lots of researchers' attention because of its goal-oriented nature. This paper studies weakly supervised sequential video understanding where the accurate time-stamp level text-video alignment is not provided. We solve this task by borrowing ideas from CLIP. Specifically, we use a transformer to aggregate frame-level features for video representation and use a pre-trained text encoder to encode the texts corresponding to each action and the whole video, respectively. To model the correspondence between text and video, we propose a multiple granularity loss, where the video-paragraph contrastive loss enforces matching between the whole video and the complete script, and a fine-grained frame-sentence contrastive loss enforces the matching between each action and its description. As the frame-sentence correspondence is not available, we propose to use the fact that video actions happen sequentially in the temporal domain to generate pseudo frame-sentence correspondence and supervise the network training with the pseudo labels. Extensive experiments on video sequence verification and textto-video matching show that our method outperforms baselines by a large margin, which validates the effectiveness of our proposed approach. Code is available at https: //github.com/svip-lab/WeakSVR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual ModalitiesA. J. Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo 等CVPR 2024 · 被引用 12 次
- Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisTianyao He, Huabin Liu, Yuxi Li, Xiao Ma 等AAAI 2024 · 被引用 8 次
- Enhanced Motion-Text Alignment for Image-to-Video Transfer LearningWei Zhang, Chaoqun Wan, Tongliang Liu, Xinmei Tian 等CVPR 2024 · 被引用 8 次
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min 等CVPR 2024 · 被引用 5 次
- De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningChong Liu, Xun Jiang, Fumin Shen, Lei Zhu 等AAAI 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 被引用 5 次
- TempCLR: Temporal Alignment Representation with Contrastive LearningYuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen 等ICLR 2023
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly DetectionZhiwei Yang, Jing Liu, Peng WuCVPR 2024 · 被引用 55 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang 等AAAI 2024 · 被引用 3 次
