Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
Min Yang, Huan Gao, Ping Guo, Limin Wang
Abstract
Vision Transformer (ViT) has shown high potential in video recognition, owing to its flexible design, adaptable self-attention mechanisms, and the efficacy of masked pretraining. Yet, it remains unclear how to adapt these pretrained short-term ViTs for temporal action detection (TAD) in untrimmed videos. The existing works treat them as off-the-shelf feature extractors for each short-trimmed snippet without capturing the fine-grained relation among different snippets in a broader temporal context. To mitigate this issue, this paper focuses on designing a new mechanism for adapting these pre-trained ViT models as a unified long-form video transformer to fully unleash its modeling power in capturing inter-snippet relation, while still keeping low computation overhead and memory consumption for efficient TAD. To this end, we design effective crosssnippet propagation modules to gradually exchange short-term video information among different snippets from two levels. For inner-backbone information propagation, we introduce a cross-snippet propagation strategy to enable multi-snippet temporal feature interaction inside the backbone. For post-backbone information propagation, we propose temporal transformer layers for further clip-level modeling. With the plain ViT-B pre-trained with VideoMAE, our end-to-end temporal action detector (ViT-TAD) yields a very competitive performance to previous temporal action detectors, riching up to 69.5 average mAP on THUMOS14, 37.40 average mAP on ActivityNet-1.3 and 17.20 average mAP on FineAction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He et al.NeurIPS 2024 · 16 citations
- Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial AdaptationTanay Agrawal, Abid Ali, Antitza Dantcheva, François BrémondICCV 2025 · 3 citations
- MobileViCLIP: An Efficient Video-Text Model for Mobile DevicesMin Yang, Zihan Jia, Zhilin Dai, Sheng Guo et al.ICCV 2025
- DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection TransformerHo-Joong Kim, Yearang Lee, Jung-Ho Hong, Seong-Whan LeeCVPR 2025
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Dual-Path Adaptation from Image to Video TransformersJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2023
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 10 citations
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
