Colar: Effective and Efficient Online Action Detection by Consulting Exemplars
Le Yang, Junwei Han, Dingwen Zhang
Abstract
Online action detection has attracted increasing research interests in recent years. Current works model historical dependencies and anticipate the future to perceive the action evolution within a video segment and improve the detection accuracy. However, the existing paradigm ignores category-level modeling and does not pay sufficient attention to efficiency. Considering a category, its representative frames exhibit various characteristics. Thus, the category-level modeling can provide complimentary guidance to the temporal dependencies modeling. This paper develops an effective exemplar-consultation mechanism that first measures the similarity between a frame and exemplary frames, and then aggregates exemplary features based on the similarity weights. This is also an efficient mechanism, as both similarity measurement and feature aggregation require limited computations. Based on the exemplar-consultation mechanism, the long-term dependencies can be captured by regarding historical frames as exemplars, while the category-level modeling can be achieved by regarding representative frames from a category as exemplars. Due to the complementarity from the categorylevel modeling, our method employs a lightweight architecture but achieves new high performance on three benchmarks. In addition, using a spatio-temporal network to tackle video frames, our method makes a good trade-off between effectiveness and efficiency. Code is available at https://github.com/VividLe/Online-Action-Detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Unified Transformer Tracker for Object TrackingFan Ma, Mike Zheng Shou, Linchao Zhu, Haoqi Fan et al.CVPR 2022 · 121 citations
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang et al.ICCV 2023 · 72 citations
- Robust Region Feature Synthesizer for Zero-Shot Object DetectionPeiliang Huang, Junwei Han, De Cheng, Dingwen ZhangCVPR 2022 · 50 citations
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He et al.NeurIPS 2024 · 16 citations
- E2E-LOAD: End-to-End Long-form Online Action DetectionShuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang et al.ICCV 2023 · 12 citations
Builds on11
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
- Enriching Local and Global Contexts for Temporal Action LocalizationZixin Zhu, Wei Tang, Le Wang, Nanning Zheng et al.ICCV 2021 · 134 citations
Related papers
- Selective Dependency Aggregation for Action ClassificationYi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei et al.ACM MM 2021 · 31 citations
- CAA: Candidate-Aware Aggregation for Temporal Action DetectionYifan Ren, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 3 citations
- Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action DetectionSu Yan, Jiahua Li, Kun Wei, Cheng DengAAAI 2026
- TS-ILM: Class Incremental Learning for Online Action DetectionXiaochen Li, Jian Cheng, Ziying Xia, Zichong Chen et al.ACM MM 2024 · 2 citations
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
