TxVAD: Improved Video Action Detection by Transformers
Zhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang, Gang Hua
Abstract
Video action detection aims to localize persons in both space and time from video sequences and recognize their actions. Most existing methods are composed of many specialized components, e.g., pretrained person/object detectors, region proposal networks (RPN), memory banks, and so on. This paper proposes a conceptually simple paradigm for video action detection using Transformers, which effectively removes the need for specialized components and achieves superior performance. Our proposed Transformer-based Video Action Detector (TxVAD) utilizes two Transformers to capture scene context information and long-range spatio-temporal context information, for person localization and action classification, respectively. Through extensive experiments on four public datasets, AVA, AVA-Kinetics, JHMDB-21, and UCF101-24, we show that our conceptually simple paradigm has achieved state-of-the-art performance for video action detection task, without using pre-trained person/object detectors, RPN, or memory bank.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e160b1d0-e262-48bc-93e4-e9870f1c19c0Cited by top-tier papers2
- LGViT: Dynamic Early Exiting for Accelerating Vision TransformerGuanyu Xu, Jiawei Hao, Li Shen, Han Hu et al.ACM MM 2023 · 33 citations
- SOAR: Scene-debiasing Open-set Action RecognitionYuanhao Zhai, Ziyi Liu, Zhenyu Wu, Yi Wu et al.ICCV 2023 · 15 citations
Related papers
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen et al.CVPR 2022 · 77 citations
- End-to-End Spatio-Temporal Action Localisation with Video TransformersAlexey A. Gritsenko, Xuehan Xiong, Josip Djolonga, Mostafa Dehghani et al.CVPR 2024
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 220 citations
