An Empirical Study of End-to-End Temporal Action Detection
Xiaolong Liu, Song Bai, Xiang Bai
Abstract
Temporal action detection (TAD) is an important yet challenging task in video understanding. It aims to simultaneously predict the semantic label and the temporal interval of every action instance in an untrimmed video. Rather than end-to-end learning, most existing methods adopt a head-only learning paradigm, where the video encoder is pre-trained for action classification, and only the detection head upon the encoder is optimized for TAD. The effect of end-to-end learning is not systematically evaluated. Besides, there lacks an in-depth study on the efficiencyaccuracy trade-off in end-to-end TAD. In this paper, we present an empirical study of end-to-end temporal action detection. We validate the advantage of end-to-end learning over head-only learning and observe up to 11% performance improvement. Besides, we study the effects of multiple design choices that affect the TAD performance and speed, including detection head, video encoder, and resolution of input videos. Based on the findings, we build a mid-resolution baseline detector, which achieves the stateof-the-art performance of end-to-end methods while running more than 4× faster. We hope that this paper can serve as a guide for end-to-end learning and inspire future research in this field. Code and models are available at https://github.com/xlliu7/E2E-TAD . * Corresponding author 1 Also known as temporal action localization (TAL).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a0b50f9-462a-425d-a809-b0274016ac4cCited by top-tier papers19
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 50 citations
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang et al.NeurIPS 2022 · 41 citations
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 35 citations
- Self-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Jae-Pil HeoICCV 2023 · 33 citations
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
Builds on16
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
Related papers
- Ex Pede Herculem, Predicting Global Actionness Curve from Local ClipsXu Chen, Yang Li, Yahong Han, Jialie ShenACM MM 2025
- Learning Salient Boundary Feature for Anchor-free Temporal Action LocalizationChuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang et al.CVPR 2021
- TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate ExpressionHo-Joong Kim, Jung-Ho Hong, Heejo Kong, Seong-Whan LeeCVPR 2024
- Post-Processing Temporal Action DetectionSauradip Nag, Xiatian Zhu, Yi-Zhe Song, Tao XiangCVPR 2023
- RefineTAD: Learning Proposal-free Refinement for Temporal Action DetectionYue Feng, Zhengye Zhang, Rong Quan, Limin Wang et al.ACM MM 2023 · 8 citations
