Watch Only Once: An End-to-End Video Action Detection Framework
Shoufa Chen, Peize Sun, Enze Xie, Chongjian Ge, Jiannan Wu, Lan Ma, Jiajun Shen, Ping Luo
摘要
We propose an end-to-end pipeline, named Watch Once Only (WOO), for video action detection. Current methods either decouple video action detection task into separated stages of actor localization and action classification or train two separated models within one stage. In contrast, our approach solves the actor localization and action classification simultaneously in a unified network. The whole pipeline is significantly simplified by unifying the backbone network and eliminating many hand-crafted components. WOO takes a unified video backbone to simultaneously extract features for actor location and action classification. In addition, we introduce spatial-temporal action embeddings into our framework and design a spatial-temporal fusion module to obtain more discriminative features with richer information, which further boosts the action classification performance. Extensive experiments on AVA and JHMDB datasets show that WOO achieves state-of-the-art performance, while still reduces up to 16.7% GFLOPs compared with existing methods. We hope our work can inspire re-thinking the convention of action detection and serve as a solid baseline for end-to-end action detection. Code is available at https://github.com/ShoufaChen/WOO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu 等ICCV 2023 · 被引用 31 次
- Multiscale Vision Transformers Meet Bipartite Matching for Efficient Single-Stage Action LocalizationIoanna Ntinou, Enrique Sanchez, Georgios TzimiropoulosCVPR 2024 · 被引用 8 次
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 被引用 5 次
它引用的顶会 Paper5
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Actor-Context-Actor Relation Network for Spatio-Temporal Action LocalizationJunting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu 等CVPR 2021
- Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsPeize Sun, Rufeng Zhang, Yi Jiang, Tao Kong 等CVPR 2021
- EfficientDet: Scalable and Efficient Object DetectionMingxing Tan, Ruoming Pang, Quoc V. LeCVPR 2020
相关 Paper
- STMixer: A One-Stage Sparse Action DetectorTao Wu, Mengqi Cao, Ziteng Gao, Gangshan Wu 等CVPR 2023
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang 等ACM MM 2022 · 被引用 5 次
- Mamba Only Glances Once (MOGO): A Lightweight Framework for Efficient Video Action DetectionYunqing Liu, Nan Zhang, Fangjun Wang, Kengo Murata 等NeurIPS 2025 · 被引用 1 次
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 被引用 73 次
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 被引用 31 次
