Self-Feedback DETR for Temporal Action Detection
Jihwan Kim, Miso Lee, Jae-Pil Heo
摘要
Temporal Action Detection (TAD) is challenging but fundamental for real-world video applications. Recently, DETR-based models have been devised for TAD but have not performed well yet. In this paper, we point out the problem in the self-attention of DETR for TAD; the attention modules focus on a few key elements, called temporal collapse problem. It degrades the capability of the encoder and decoder since their self-attention modules play no role. To solve the problem, we propose a novel framework, Self-DETR, which utilizes cross-attention maps of the decoder to reactivate self-attention modules. We recover the relationship between encoder features by simple matrix multi-plication of the cross-attention map and its transpose. Likewise, we also get the information within decoder queries. By guiding collapsed self-attention maps with the guidance map calculated, we settle down the temporal collapse of self-attention modules in the encoder and decoder. Our extensive experiments demonstrate that Self-DETR resolves the temporal collapse problem by keeping high diversity of attention over all layers. Moreover, it is validated that our simple framework achieves a new state-of-the-art performance on THUMOS14 and outperforms all the DETR-based approaches on ActivityNet-v1.3.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Dual DETRs for Multi-Label Temporal Action DetectionYuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu 等CVPR 2024 · 被引用 25 次
- Prediction-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Cheol-Ho Cho, Jihyun Lee 等AAAI 2025 · 被引用 8 次
- Activating Self-Attention for Multi-Scene Absolute Pose RegressionMiso Lee, Jihwan Kim, Jae-Pil HeoNeurIPS 2024 · 被引用 4 次
- Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationHaoyu Tang, Tianyuan Liang, Han Jiang, Xuesong Liu 等AAAI 2026
- Beyond Caption-Based Queries in Video Moment RetrievalDavid Pujol-Perich, Albert Clapés, Dima Damen, Sergio Escalera 等CVPR 2026
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
相关 Paper
- DiffTAD: Temporal Action Detection with Proposal Denoising DiffusionSauradip Nag, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song 等ICCV 2023 · 被引用 34 次
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 被引用 220 次
- Class Semantics-based Attention for Action DetectionDeepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li 等ICCV 2021 · 被引用 77 次
- Sim-DETR: Unlock DETR for Temporal Sentence GroundingJiajin Tang, Zhengxuan Wei, Yuchen Zhu, Cheng Shi 等ICCV 2025 · 被引用 3 次
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu 等ACM MM 2021 · 被引用 106 次
