Beyond Short Clips: End-to-End Video-Level Learning With Collaborative Memories
Xitong Yang, Haoqi Fan, Lorenzo Torresani, Larry S. Davis, Heng Wang
摘要
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We argue that a single clip may not have enough temporal coverage to exhibit the label to recognize, since video datasets are often weakly labeled with categorical information but without dense temporal annotations. Furthermore, optimizing the model over brief clips impedes its ability to learn long-term temporal dependencies. To overcome these limitations, we introduce a collaborative memory mechanism that encodes information across multiple sampled clips of a video at each training iteration. This enables the learning of long-range dependencies beyond a single clip. We explore different design choices for the collaborative memory to ease the optimization difficulties. Our proposed framework is end-to-end trainable and significantly improves the accuracy of video classification at a negligible computational overhead. Through extensive experiments, we demonstrate that our framework generalizes to different video architectures and tasks, outperforming the state of the art on both action recognition (e.g., Charades, AVA v2.1 & v2.2).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical ConsistencyZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu 等CVPR 2022 · 被引用 11 次
- Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action DetectionZhe Luo, Weina Fu, Shuai Liu, Saeed Anwar 等ACM MM 2024 · 被引用 4 次
- Relational Space-Time Query in Long-Form VideosXitong Yang, Fu-Jen Chu, Matt Feiszli, Raghav Goyal 等CVPR 2023
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionSanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan 等ICCV 2025
它引用的顶会 Paper10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
相关 Paper
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
- Multi-Label Activity Recognition Using Activity-Specific Features and Activity CorrelationsYanyi Zhang, Xinyu Li, Ivan MarsicCVPR 2021
- Streaming Dense Video CaptioningXingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan 等CVPR 2024 · 被引用 33 次
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 被引用 27 次
- PGT: A Progressive Method for Training Models on Long VideosBo Pang, Gao Peng, Yizhuo Li, Cewu LuCVPR 2021
