DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
Ho-Joong Kim, Yearang Lee, Jung-Ho Hong, Seong-Whan Lee
Abstract
In this paper, we examine a key limitation in query-based detectors for temporal action detection (TAD), which arises from their direct adaptation of originally designed architectures for object detection. Despite the effectiveness of the existing models, they struggle to fully address the unique challenges of TAD, such as the redundancy in multiscale features and the limited ability to capture sufficient temporal context. To address these issues, we propose a multi-dilated gated encoder and central-adjacent region integrated decoder for temporal action detection transformer (DiGIT). Our approach replaces the existing encoder that consists of multi-scale deformable attention and feedforward network with our multi-dilated gated encoder. Our proposed encoder reduces the redundant information caused by multi-level features while maintaining the ability to capture fine-grained and long-range temporal information. Furthermore, we introduce a central-adjacent region integrated decoder that leverages a more comprehensive sampling strategy for deformable cross-attention to capture the essential information. Extensive experiments demonstrate that DiGIT achieves state-of-the-art performance on THUMOS14, ActivityNet v1.3, and HACS-Segment. Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosPeijun Bao, Anwei Luo, Gang Pan, Alex C. Kot et al.CVPR 2026 · 2 citations
- Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action LocalizationJiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni et al.ACL 2026 · 1 citation
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang et al.SIGIR 2026
- Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Xiaohua Chen et al.CVPR 2026
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate ExpressionHo-Joong Kim, Jung-Ho Hong, Heejo Kong, Seong-Whan LeeCVPR 2024
- Prediction-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Cheol-Ho Cho, Jihyun Lee et al.AAAI 2025 · 8 citations
- DiffTAD: Temporal Action Detection with Proposal Denoising DiffusionSauradip Nag, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song et al.ICCV 2023 · 34 citations
- Self-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Jae-Pil HeoICCV 2023 · 33 citations
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy et al.CVPR 2026 · 5 citations
