Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts
Kunpeng Li, Zizhao Zhang, Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu, Tomas Pfister
Abstract
Learning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this issue, we introduce a new method for pretraining video action recognition models using queried web videos. Instead of trying to filter out, we propose to convert the potential noises in these queried videos to useful supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations. SPL is fairly simple and orthogonal to popular teacher-student self-training frameworks without extra training cost. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset to study the generalization ability. Experiments show that SPL outperforms several existing pre-training strategies using pseudolabels and the learned representations lead to competitive results when fine-tuning on HMDB-51 and UCF-101 compared with recent pre-training methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Anchor Loss: Modulating Loss Scale Based on Prediction DifficultySerim Ryou, Seong-Gyun Jeong, Pietro PeronaICCV 2019 · 46 citations
- SpeedNet: Learning the Speediness in VideosSagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri et al.CVPR 2020
Related papers
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo et al.AAAI 2020 · 7 citations
- Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise CorrectionQuan Zhang, Yuxin Qi, Xi Tang, Rui Yuan et al.AAAI 2025 · 11 citations
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua et al.ICCV 2023 · 16 citations
- WOAD: Weakly Supervised Online Action Detection in Untrimmed VideosMingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher et al.CVPR 2021
