Fine-grained Key-Value Memory Enhanced Predictor for Video Representation Learning
Xiaojie Li, Jianlong Wu, Shaowei He, Shuo Kang, Yue Yu, Liqiang Nie, Min Zhang
Abstract
Self-supervised learning methods have shown significant promise in acquiring robust spatiotemporal representations from unlabeled videos. In this work, we address three critical limitations in existing self-supervised video representation learning: 1) insufficient utilization of contextual information and lifelong memory, 2) lack of fine-grained visual concept alignment, and 3) neglect of the feature distribution gap between encoders. To overcome these limitations, we propose a novel memory-enhanced predictor that leverages key-value memory networks with separate memories for the online and target encoders. This design enables the effective storage and retrieval of contextual knowledge, facilitating informed predictions and enhancing overall performance. Additionally, we introduce a visual concept alignment module that ensures fine-grained alignment of shared semantic information across segments of the same video. By employing coupled dictionary learning, we effectively decouple visual concepts, enriching the semantic representation stored in the memory networks. Our proposed approach is extensively evaluated on widely recognized benchmarks for action recognition and retrieval tasks, demonstrating its superiority in learning generalized video representations with significantly improved performance compared to existing state-of-the-art self-supervised learning methods. Code is released at https://github.com/xiaojieli0903/FGKVMemPred_video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42b8465f-aaff-4c77-ac54-35a52aef15e3Cited by top-tier papers5
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen et al.NeurIPS 2024 · 104 citations
- Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic ManipulationQi Lv, Hao Li, Xiang Deng, Rui Shao et al.CVPR 2025
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan et al.CVPR 2025
- Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned PolicyZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen et al.CVPR 2025
- CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuningYibo Yang, Xiaojie Li, Zhongzhu Zhou, Shuaiwen Song et al.NeurIPS 2024
Builds on25
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
- MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation LearningBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ACM MM 2022 · 6 citations
- Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationAli Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi et al.ICCV 2021 · 65 citations
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen et al.AAAI 2022 · 35 citations
