Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts
Kunpeng Li, Zizhao Zhang, Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu, Tomas Pfister
摘要
Learning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this issue, we introduce a new method for pretraining video action recognition models using queried web videos. Instead of trying to filter out, we propose to convert the potential noises in these queried videos to useful supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations. SPL is fairly simple and orthogonal to popular teacher-student self-training frameworks without extra training cost. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset to study the generalization ability. Experiments show that SPL outperforms several existing pre-training strategies using pseudolabels and the learned representations lead to competitive results when fine-tuning on HMDB-51 and UCF-101 compared with recent pre-training methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- Anchor Loss: Modulating Loss Scale Based on Prediction DifficultySerim Ryou, Seong-Gyun Jeong, Pietro PeronaICCV 2019 · 被引用 46 次
- SpeedNet: Learning the Speediness in VideosSagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri 等CVPR 2020
相关 Paper
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah 等CVPR 2024 · 被引用 3 次
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo 等AAAI 2020 · 被引用 7 次
- Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise CorrectionQuan Zhang, Yuxin Qi, Xi Tang, Rui Yuan 等AAAI 2025 · 被引用 11 次
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua 等ICCV 2023 · 被引用 16 次
- WOAD: Weakly Supervised Online Action Detection in Untrimmed VideosMingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher 等CVPR 2021
