Movement Enhancement toward Multi-Scale Video Feature Representation for Temporal Action Detection
Zixuan Zhao, Dongqi Wang, Xu Zhao
摘要
Boundary localization is a challenging problem in Temporal Action Detection (TAD), in which there are two main issues. First, the submergence of movement feature, i.e. the movement information in a snippet is covered by the scene information. Second, the scale of action, that is, the proportion of action segments in the entire video, is considerably variable. In this work, we first design a Movement Enhance Module (MEM) to highlight movement feature for better action location, and then, we propose a Scale Feature Pyramid Network (SFPN) to detect multi-scale actions in videos. For Movement Enhance Module, firstly, Movement Feature Extractor (MFE) is designed to get the movement feature. Secondly, we propose a Multi-Relation Enhance Module (MREM) to grasp valuable information correlation both locally and temporally. For Scale Feature Pyramid Network, we design a U-Shape Module to model different scale actions, moreover, we design the training and inference strategy of different scales, ensuring that each pyramid layer is only responsible for actions at a specific scale. These two innovations are integrated as the Movement Enhance Network (MENet), and extensive experiments conducted on two challenging benchmarks demonstrate its effectiveness. MENet outperforms other representative TAD methods on ActivityNet-1.3 and THUMOS-14.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 被引用 35 次
- Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationHaoyu Tang, Tianyuan Liang, Han Jiang, Xuesong Liu 等AAAI 2026
- TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate ExpressionHo-Joong Kim, Jung-Ho Hong, Heejo Kong, Seong-Whan LeeCVPR 2024
它引用的顶会 Paper9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai 等AAAI 2020 · 被引用 226 次
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 被引用 220 次
- BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal GenerationHaisheng Su, Weihao Gan, Wei Wu, Yu Qiao 等AAAI 2021 · 被引用 143 次
相关 Paper
- Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid NetworkJialin Gao, Zhixiang Shi, Guanshuo Wang, Jiani Li 等AAAI 2020 · 被引用 78 次
- TriDet: Temporal Action Detection with Relative Boundary ModelingDingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma 等CVPR 2023
- Progressive Boundary Refinement Network for Temporal Action DetectionQinying Liu, Zilei WangAAAI 2020 · 被引用 156 次
- Pixels, Regions, and Objects: Multiple Enhancement for Salient Object DetectionYi Wang, Ruili Wang, Xin Fan, Tianzhu Wang 等CVPR 2023
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li 等AAAI 2025 · 被引用 3 次
