Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization
Chen Zhao, Shuming Liu, Karttikeya Mangalam, Bernard Ghanem
摘要
Temporal action localization (TAL) requires long-form reasoning to predict actions of various durations and complex content. Given limited GPU memory, training TAL end to end (i.e., from videos to predictions) on long videos is a significant challenge. Most methods can only train on pre-extracted features without optimizing them for the localization problem, consequently limiting localization performance. In this work, to extend the potential in TAL networks, we propose a novel end-to-end method Re 2 TAL, which rewires pretrained video backbones for reversible TAL. Re 2 TAL builds a backbone with reversible modules, where the input can be recovered from the output such that the bulky intermediate activations can be cleared from memory during training. Instead of designing one single type of reversible module, we propose a network rewiring mechanism, to transform any module with a residual connection to a reversible module without changing any parameters. This provides two benefits: (1) a large variety of reversible networks are easily obtained from existing and even future model designs, and (2) the reversible models require much less training effort as they reuse the pre-trained parameters of their original non-reversible versions. Re 2 TAL, only using the RGB modality, reaches 37.01% average mAP on ActivityNet-v1.3, a new state-ofthe-art record, and mAP 64.9% at tIoU=0.5 on THUMOS-14, outperforming all other RGB-only methods. Code is available at https://github.com/coolbay/Re2TAL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 被引用 35 次
- EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesJinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao 等ICCV 2023 · 被引用 26 次
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang 等ICCV 2025 · 被引用 21 次
- Adapting Short-Term Transformers for Action Detection in Untrimmed VideosMin Yang, Huan Gao, Ping Guo, Limin WangCVPR 2024 · 被引用 14 次
- Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial AdaptationTanay Agrawal, Abid Ali, Antitza Dantcheva, François BrémondICCV 2025 · 被引用 3 次
它引用的顶会 Paper27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- Low-Fidelity Video Encoder Optimization for Temporal Action LocalizationMengmeng Xu, Juan-Manuel Pérez-Rúa, Xiatian Zhu, Bernard Ghanem 等NeurIPS 2021 · 被引用 30 次
- Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationFa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan 等ACM MM 2021 · 被引用 104 次
- A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationZan Gao, Xinglei Cui, Yibo Zhao, Tao Zhuo 等ACM MM 2023 · 被引用 2 次
- Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient FinetuningChen Zhao, Shuming Liu, Karttikeya Mangalam, Guocheng Qian 等CVPR 2024
- Learning Generalized Representations for Open-Set Temporal Action LocalizationJunshan Hu, Liansheng Zhuang, Weisong Dong, Shiming Ge 等ACM MM 2023 · 被引用 4 次
