Multi-dataset Training of Transformers for Robust Action Recognition
Junwei Liang, Enwei Zhang, Jun Zhang, Chunhua Shen
Abstract
We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action recognition in the past decade, it remains challenging yet valuable how to train a single model that can perform well across multiple datasets. Here, we propose a novel multidataset training paradigm, MultiTrain, with the design of two new loss terms, namely informative loss and projection loss, aiming to learn robust representations for action recognition. In particular, the informative loss maximizes the expressiveness of the feature embedding while the projection loss for each dataset mines the intrinsic relations between classes across datasets. We verify the effectiveness of our method on five challenging datasets, Kinetics-400, Kinetics-700, Moments-in-Time, Activitynet and Something-something-v2 datasets. Extensive experimental results show that our method can consistently improve state-of-the-art performance. Code and models are available at https://github.com/JunweiLiang/MultiTrain
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance SegmentationRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.NeurIPS 2023 · 3 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- DVANet: Disentangling View and Action Features for Multi-View Action RecognitionNyle Siddiqui, Praveen Tirupattur, Mubarak ShahAAAI 2024 · 39 citations
- Can an Image Classifier Suffice For Action Recognition?Quanfu Fan, Chun-Fu Chen, Rameswar PandaICLR 2022 · 39 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 40 citations
- BEVT: BERT Pretraining of Video TransformersRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2022 · 200 citations
