Relaxed Transformer Decoders for Direct Action Proposal Generation
Jing Tan, Jiaqi Tang, Limin Wang, Gangshan Wu
Abstract
Temporal action proposal generation is an important and challenging task in video understanding, which aims at detecting all temporal segments containing action in-stances of interest. The existing proposal generation approaches are generally based on pre-defined anchor windows or heuristic bottom-up boundary matching strategies. This paper presents a simple and efficient framework (RTD-Net) for direct action proposal generation, by re-purposing a Transformer-alike architecture. To tackle the essential visual difference between time and space, we make three important improvements over the original transformer detection framework (DETR). First, to deal with slowness prior in videos, we replace the original Transformer en-coder with a boundary attentive module to better capture long-range temporal information. Second, due to the ambiguous temporal boundary and relatively sparse annotations, we present a relaxed matching scheme to relieve the strict criteria of single assignment to each groundtruth. Finally, we devise a three-branch head to further improve the proposal confidence estimation by explicitly predicting its completeness. Extensive experiments on THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of RTD-Net, on both tasks of temporal action proposal generation and temporal action detection. Moreover, due to its simplicity in design, our framework is more efficient than previous proposal generation methods, without non-maximum suppression post-processing. The code and models are made available at https://github.com/MCG-NJU/RTD-Action.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 716e85ad-3886-4535-8034-aead679808c8Cited by top-tier papers52
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo et al.CVPR 2022 · 93 citations
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
- DCAN: Improving Temporal Action Detection via Dual Context AggregationGuo Chen, Yin-Dong Zheng, Limin Wang, Tong LuAAAI 2022 · 86 citations
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 80 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
Related papers
- DiffTAD: Temporal Action Detection with Proposal Denoising DiffusionSauradip Nag, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song et al.ICCV 2023 · 34 citations
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
- Prediction-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Cheol-Ho Cho, Jihyun Lee et al.AAAI 2025 · 8 citations
- Self-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Jae-Pil HeoICCV 2023 · 33 citations
