Deep Concept-wise Temporal Convolutional Networks for Action Localization
Xin Li, Tianwei Lin, Xiao Liu, Wangmeng Zuo, Chao Li, Xiang Long, Dongliang He, Fu Li, Shilei Wen, Chuang Gan
Abstract
Existing action localization approaches adopt shallow temporal convolutional networks (i.e., TCN) on 1D feature map extracted from video frames. In this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which generally are highly abstract and can be regarded as latent concepts, are excessively recombined in temporal convolution. To address this issue, we introduce a novel concept-wise temporal convolutional network (C-TCN) as an alternative to TCN for training deeper action localization networks. To address this issue, we introduce a novel concept-wise temporal convolution (CTC) layer as an alternative to conventional temporal convolution layer for training deeper action localization networks. Instead of recombining latent concepts, CTC layer deploys a number of temporal filters to each concept separately with shared filter parameters across concepts. Thus can capture common temporal patterns of different concepts and significantly enrich representation ability. Via stacking CTC layers, we proposed a deep concept-wise temporal convolutional network (C-TCN), which boosts the state-of-the-art action localization performance on THUMOS'14 from 42.8 to 52.1 in terms of mAP(%), achieving a relative improvement of 21.7%. Favorable result is also obtained on ActivityNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83b8bd16-ee07-425e-ae93-bafe1e3c16ecCited by top-tier papers4
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- SVIP: Sequence VerIfication for Procedures in VideosYicheng Qian, Weixin Luo, Dongze Lian, Xu Tang et al.CVPR 2022 · 19 citations
- Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target DetectionHouzhang Fang, Shukai Guo, Qiuhuan Chen, Yi Chang et al.AAAI 2026
- Multi-Stage Aggregated Transformer Network for Temporal Language Localization in VideosMingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji et al.CVPR 2021
Related papers
- A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationZan Gao, Xinglei Cui, Yibo Zhao, Tao Zhuo et al.ACM MM 2023 · 2 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Weakly Supervised Action Selection Learning in VideoJunwei Ma, Satya Krishna Gorti, Maksims Volkovs, Guangwei YuCVPR 2021
- Class Semantics-based Attention for Action DetectionDeepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li et al.ICCV 2021 · 77 citations
- ACGNet: Action Complement Graph Network for Weakly-Supervised Temporal Action LocalizationZichen Yang, Jie Qin, Di HuangAAAI 2022 · 72 citations
