TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
Ho-Joong Kim, Jung-Ho Hong, Heejo Kong, Seong-Whan Lee
摘要
In this paper, we investigate that the normalized coordinate expression is a key factor as reliance on handcrafted components in query-based detectors for temporal action detection (TAD). Despite significant advancements towards an end-to-end framework in object detection, query-based detectors have been limited in achieving full end-to-end modeling in TAD. To address this issue, we propose TE-TAD, a full end-to-end temporal action detection transformer that integrates time-aligned coordinate expression. We reformulate coordinate expression utilizing actual timeline values, ensuring length-invariant representations from the extremely diverse video duration environment. Furthermore, our proposed adaptive query selection dynamically adjusts the number of queries based on video length, providing a suitable solution for varying video durations compared to a fixed query set. Our approach not only simplifies the TAD process by eliminating the need for hand-crafted components but also significantly improves the performance of query-based detectors. Our TE-TAD outperforms the previous query-based detectors and achieves competitive performance compared to stateof-the-art methods on popular benchmark datasets. Code is
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action DetectionYearang Lee, Ho-Joong Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 12 次
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionYupeng Hu, Changxing Ding, Chang Sun, Shaoli Huang 等ICCV 2025 · 被引用 1 次
- DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery LocalizationXiaodong Zhu, Suting Wang, Yuanming Zheng, Junqi Yang 等AAAI 2026
- Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationHaoyu Tang, Tianyuan Liang, Han Jiang, Xuesong Liu 等AAAI 2026
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang 等SIGIR 2026
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection TransformerHo-Joong Kim, Yearang Lee, Jung-Ho Hong, Seong-Whan LeeCVPR 2025
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu 等ACM MM 2021 · 被引用 106 次
- An Empirical Study of End-to-End Temporal Action DetectionXiaolong Liu, Song Bai, Xiang BaiCVPR 2022 · 被引用 72 次
- Ex Pede Herculem, Predicting Global Actionness Curve from Local ClipsXu Chen, Yang Li, Yahong Han, Jialie ShenACM MM 2025
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang 等ACM MM 2022 · 被引用 5 次
