Beyond Short Clips: End-to-End Video-Level Learning With Collaborative Memories
Xitong Yang, Haoqi Fan, Lorenzo Torresani, Larry S. Davis, Heng Wang
Abstract
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We argue that a single clip may not have enough temporal coverage to exhibit the label to recognize, since video datasets are often weakly labeled with categorical information but without dense temporal annotations. Furthermore, optimizing the model over brief clips impedes its ability to learn long-term temporal dependencies. To overcome these limitations, we introduce a collaborative memory mechanism that encodes information across multiple sampled clips of a video at each training iteration. This enables the learning of long-range dependencies beyond a single clip. We explore different design choices for the collaborative memory to ease the optimization difficulties. Our proposed framework is end-to-end trainable and significantly improves the accuracy of video classification at a negligible computational overhead. Through extensive experiments, we demonstrate that our framework generalizes to different video architectures and tasks, outperforming the state of the art on both action recognition (e.g., Charades, AVA v2.1 & v2.2).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed845615-cfc1-4def-8ef2-39d8fa109edaCited by top-tier papers7
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
- Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical ConsistencyZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu et al.CVPR 2022 · 11 citations
- Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action DetectionZhe Luo, Weina Fu, Shuai Liu, Saeed Anwar et al.ACM MM 2024 · 4 citations
- Relational Space-Time Query in Long-Form VideosXitong Yang, Fu-Jen Chu, Matt Feiszli, Raghav Goyal et al.CVPR 2023
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionSanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan et al.ICCV 2025
Builds on10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
Related papers
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan et al.CVPR 2022 · 158 citations
- Multi-Label Activity Recognition Using Activity-Specific Features and Activity CorrelationsYanyi Zhang, Xinyu Li, Ivan MarsicCVPR 2021
- Streaming Dense Video CaptioningXingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan et al.CVPR 2024 · 33 citations
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 27 citations
- PGT: A Progressive Method for Training Models on Long VideosBo Pang, Gao Peng, Yizhuo Li, Cewu LuCVPR 2021
