Low-Fidelity Video Encoder Optimization for Temporal Action Localization
Mengmeng Xu, Juan-Manuel Pérez-Rúa, Xiatian Zhu, Bernard Ghanem, Brais Martínez
Abstract
Most existing temporal action localization (TAL) methods rely on a transfer learning pipeline, first optimizing a video encoder on a large action classification dataset (i.e., source domain), followed by freezing the encoder and training a TAL head on the action localization dataset (i.e., target domain). This results in a task discrepancy problem for the video encoder -trained for action classification, but used for TAL. Intuitively, joint optimization with both the video encoder and TAL head is an obvious solution to this discrepancy. However, this is not operable for TAL subject to the GPU memory constraints, due to the prohibitive computational cost in processing long untrimmed videos. In this paper, we resolve this challenge by introducing a novel low-fidelity (LoFi) video encoder optimization method. Instead of always using the full training configurations in TAL learning, we propose to reduce the mini-batch composition in terms of temporal, spatial or spatio-temporal resolution so that jointly optimizing the video encoder and TAL head becomes operable under the same memory conditions of a mid-range hardware budget. Crucially, this enables the gradients to flow backwards through the video encoder conditioned on a TAL supervision loss, favourably solving the task discrepancy problem and providing more effective feature representations. Extensive experiments show that the proposed LoFi optimization approach can significantly enhance the performance of existing TAL methods. Encouragingly, even with a lightweight ResNet18 based video encoder in a single RGB stream, our method surpasses two-stream (RGB + optical flow) ResNet50 based alternatives, often by a good margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb238041-7e99-4123-9fb8-ea458b6292d6Cited by top-tier papers9
- An Empirical Study of End-to-End Temporal Action DetectionXiaolong Liu, Song Bai, Xiang BaiCVPR 2022 · 72 citations
- Action Sensitivity Learning for Temporal Action LocalizationJiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng et al.ICCV 2023 · 44 citations
- Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action LocalizationMaodong Li, Chao Zheng, Jian Wang, Bing LiAAAI 2025 · 2 citations
- Soft-Landing Strategy for Alleviating the Task Discrepancy Problem in Temporal Action Localization TasksHyolim Kang, Hanjung Kim, Joungbin An, Minsu Cho et al.CVPR 2023
- Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action LocalizationChen Zhao, Shuming Liu, Karttikeya Mangalam, Bernard GhanemCVPR 2023
Builds on22
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
Related papers
- Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionPilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee et al.CVPR 2023
- Learning Disentangled Classification and Localization Representations for Temporal Action LocalizationZixin Zhu, Le Wang, Wei Tang, Ziyi Liu et al.AAAI 2022 · 18 citations
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li et al.AAAI 2025 · 3 citations
- Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial AdaptationTanay Agrawal, Abid Ali, Antitza Dantcheva, François BrémondICCV 2025 · 3 citations
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
