Watch Only Once: An End-to-End Video Action Detection Framework
Shoufa Chen, Peize Sun, Enze Xie, Chongjian Ge, Jiannan Wu, Lan Ma, Jiajun Shen, Ping Luo
Abstract
We propose an end-to-end pipeline, named Watch Once Only (WOO), for video action detection. Current methods either decouple video action detection task into separated stages of actor localization and action classification or train two separated models within one stage. In contrast, our approach solves the actor localization and action classification simultaneously in a unified network. The whole pipeline is significantly simplified by unifying the backbone network and eliminating many hand-crafted components. WOO takes a unified video backbone to simultaneously extract features for actor location and action classification. In addition, we introduce spatial-temporal action embeddings into our framework and design a spatial-temporal fusion module to obtain more discriminative features with richer information, which further boosts the action classification performance. Extensive experiments on AVA and JHMDB datasets show that WOO achieves state-of-the-art performance, while still reduces up to 16.7% GFLOPs compared with existing methods. We hope our work can inspire re-thinking the convention of action detection and serve as a solid baseline for end-to-end action detection. Code is available at https://github.com/ShoufaChen/WOO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan et al.CVPR 2022 · 158 citations
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen et al.CVPR 2022 · 77 citations
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
- Multiscale Vision Transformers Meet Bipartite Matching for Efficient Single-Stage Action LocalizationIoanna Ntinou, Enrique Sanchez, Georgios TzimiropoulosCVPR 2024 · 8 citations
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 5 citations
Builds on5
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Actor-Context-Actor Relation Network for Spatio-Temporal Action LocalizationJunting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu et al.CVPR 2021
- Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsPeize Sun, Rufeng Zhang, Yi Jiang, Tao Kong et al.CVPR 2021
- EfficientDet: Scalable and Efficient Object DetectionMingxing Tan, Ruoming Pang, Quoc V. LeCVPR 2020
Related papers
- STMixer: A One-Stage Sparse Action DetectorTao Wu, Mengqi Cao, Ziteng Gao, Gangshan Wu et al.CVPR 2023
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- Mamba Only Glances Once (MOGO): A Lightweight Framework for Efficient Video Action DetectionYunqing Liu, Nan Zhang, Fangjun Wang, Kengo Murata et al.NeurIPS 2025 · 1 citation
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 73 citations
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 31 citations
