SVFormer: Semi-supervised Video Transformer for Action Recognition
Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang
摘要
Semi-supervised action recognition is a challenging but critical task due to the high cost of video annotations. Existing approaches mainly use convolutional neural networks, yet current revolutionary vision transformer models have been less explored. In this paper, we investigate the use of transformer models under the SSL setting for action recognition. To this end, we introduce SVFormer, which adopts a steady pseudo-labeling framework (i.e., EMA-Teacher) to cope with unlabeled video samples. While a wide range of data augmentations have been shown effective for semi-supervised image classification, they generally produce limited results for video recognition. We therefore introduce a novel augmentation strategy, Tube Token-Mix, tailored for video data where video clips are mixed via a mask with consistent masked tokens over the temporal axis. In addition, we propose a temporal warping augmentation to cover the complex temporal variation in videos, which stretches selected frames to various temporal durations in the clip. Extensive experiments on three datasets Kinetics-400, UCF-101, and HMDB-51 verify the advantage of SVFormer. In particular, SVFormer outperforms the state-of-the-art by 31.5% with fewer training epochs under the 1% labeling rate of Kinetics-400. Our method can hopefully serve as a strong benchmark and encourage future search on semi-supervised action recognition with Transformer networks. Code is released at https: //github.com/ChenHsing/SVFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 等ICML 2023 · 被引用 67 次
- Implicit Temporal Modeling with Learnable Alignment for Video RecognitionShuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng 等ICCV 2023 · 被引用 63 次
- SimDA: Simple Diffusion Adapter for Efficient Video GenerationZhen Xing, Qi Dai, Han Hu, Zuxuan Wu 等CVPR 2024 · 被引用 34 次
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 被引用 27 次
- MotionEditor: Editing Video Motion via Content-Aware DiffusionShuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu 等CVPR 2024 · 被引用 21 次
它引用的顶会 Paper37
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- Semi-supervised Vision Transformers at ScaleZhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang 等NeurIPS 2022 · 被引用 82 次
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 等CVPR 2022 · 被引用 111 次
- Can an Image Classifier Suffice For Action Recognition?Quanfu Fan, Chun-Fu Chen, Rameswar PandaICLR 2022 · 被引用 39 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Training Vision Transformers for Semi-Supervised Semantic SegmentationXinting Hu, Li Jiang, Bernt SchieleCVPR 2024
