PGT: A Progressive Method for Training Models on Long Videos
Bo Pang, Gao Peng, Yizhuo Li, Cewu Lu
Abstract
Convolutional video models have an order of magnitude larger computational complexity than their counterpart image-level models. Constrained by computational resources, there is no model or training method that can train long video sequences end-to-end. Currently, the mainstream method is to split a raw video into clips, leading to incomplete fragmentary temporal information flow. Inspired by natural language processing techniques dealing with long sentences, we propose to treat videos as serial fragments satisfying Markov property, and train it as a whole by progressively propagating information through the temporal dimension in multiple steps. This progressive training (PGT) method is able to train long videos end-to-end with limited resources and ensures the effective transmission of information. As a general and robust training method, we empirically demonstrate that it yields significant performance improvements on different models and datasets. As an illustrative example, the proposed method improves SlowOnly network by 3.7 mAP on Charades and 1.9 top-1 accuracy on Kinetics with negligible parameter and computation overhead. Code is available at: https://github.com/BoPang1996/PGT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc3d4ec3-a151-46e9-83c7-7bb4a0904e8aCited by top-tier papers3
- Highlighting Object Category Immunity for the Generalization of Human-Object Interaction DetectionXinpeng Liu, Yong-Lu Li, Cewu LuAAAI 2022 · 16 citations
- Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical ConsistencyZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu et al.CVPR 2022 · 11 citations
- Understanding Dynamic Scenes in Ego Centric 4D Point CloudsJunsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang et al.AAAI 2026 · 4 citations
Builds on13
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 109 citations
- Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNsLei Wang, Piotr Koniusz, Du HuynhICCV 2019 · 100 citations
Related papers
- A Multigrid Method for Efficiently Training Video ModelsChao-Yuan Wu, Ross B. Girshick, Kaiming He, Christoph Feichtenhofer et al.CVPR 2020
- Beyond Short Clips: End-to-End Video-Level Learning With Collaborative MemoriesXitong Yang, Haoqi Fan, Lorenzo Torresani, Larry S. Davis et al.CVPR 2021
- Video-GPT via Next Clip DiffusionShaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang et al.ICLR 2026 · 9 citations
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive LearningChanglin Li, Jiawei Zhang, Shuhao Liu, Sihao Lin et al.CVPR 2026 · 2 citations
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 37 citations
