Self-Feedback DETR for Temporal Action Detection
Jihwan Kim, Miso Lee, Jae-Pil Heo
Abstract
Temporal Action Detection (TAD) is challenging but fundamental for real-world video applications. Recently, DETR-based models have been devised for TAD but have not performed well yet. In this paper, we point out the problem in the self-attention of DETR for TAD; the attention modules focus on a few key elements, called temporal collapse problem. It degrades the capability of the encoder and decoder since their self-attention modules play no role. To solve the problem, we propose a novel framework, Self-DETR, which utilizes cross-attention maps of the decoder to reactivate self-attention modules. We recover the relationship between encoder features by simple matrix multi-plication of the cross-attention map and its transpose. Likewise, we also get the information within decoder queries. By guiding collapsed self-attention maps with the guidance map calculated, we settle down the temporal collapse of self-attention modules in the encoder and decoder. Our extensive experiments demonstrate that Self-DETR resolves the temporal collapse problem by keeping high diversity of attention over all layers. Moreover, it is validated that our simple framework achieves a new state-of-the-art performance on THUMOS14 and outperforms all the DETR-based approaches on ActivityNet-v1.3.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ad84ef0-2cc5-452f-bfcf-a3ddf888e1b0Cited by top-tier papers8
- Dual DETRs for Multi-Label Temporal Action DetectionYuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu et al.CVPR 2024 · 25 citations
- Prediction-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Cheol-Ho Cho, Jihyun Lee et al.AAAI 2025 · 8 citations
- Activating Self-Attention for Multi-Scene Absolute Pose RegressionMiso Lee, Jihwan Kim, Jae-Pil HeoNeurIPS 2024 · 4 citations
- Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationHaoyu Tang, Tianyuan Liang, Han Jiang, Xuesong Liu et al.AAAI 2026
- Beyond Caption-Based Queries in Video Moment RetrievalDavid Pujol-Perich, Albert Clapés, Dima Damen, Sergio Escalera et al.CVPR 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
Related papers
- DiffTAD: Temporal Action Detection with Proposal Denoising DiffusionSauradip Nag, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song et al.ICCV 2023 · 34 citations
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 220 citations
- Class Semantics-based Attention for Action DetectionDeepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li et al.ICCV 2021 · 77 citations
- Sim-DETR: Unlock DETR for Temporal Sentence GroundingJiajin Tang, Zhengxuan Wei, Yuchen Zhu, Cheng Shi et al.ICCV 2025 · 3 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
