Self-Supervised Video Representation Learning via Latent Time Navigation
Di Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
摘要
Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as enter' and leave' to be indistinguishable. To mitigate this limitation, we propose Latent Time Navigation (LTN), a time parameterized contrastive learning strategy that is streamlined to capture fine-grained motions. Specifically, we maximize the representation similarity between different video segments from one video, while maintaining their representations time-aware along a subspace of the latent representation code including an orthogonal basis to represent temporal changes. Our extensive experimental analysis suggests that learning video representations by LTN consistently improves performance of action classification in fine-grained and human-oriented tasks (e.g., on Toyota Smarthome dataset). In addition, we demonstrate that our proposed model, when pre-trained on Kinetics-400, generalizes well onto the unseen real world video benchmark datasets UCF101 and HMDB51, achieving state-of-the-art performance in action recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 被引用 14 次
- Just Add π! Pose Induced Video Transformers for Understanding Activities of Daily LivingDominick Reilly, Srijan DasCVPR 2024 · 被引用 14 次
- Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action RecognitionHaochen Chang, Pengfei Ren, Haoyang Zhang, Liang Xie 等ICCV 2025 · 被引用 8 次
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World ModelsZaishuo Xia, Yukuan Lu, Xinyi Li, Yifan Xu 等CVPR 2026 · 被引用 3 次
- MoVie: Broaden Your Views with Human Motion for Action DetectionDi Yang, Mahmoud Ali, Xuanlong Yu, Xi Shen 等CVPR 2026
它引用的顶会 Paper22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 被引用 219 次
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo 等ICCV 2019 · 被引用 182 次
相关 Paper
- Contextualized Spatio-Temporal Contrastive Learning with Self-SupervisionLiangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong 等CVPR 2022 · 被引用 24 次
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 被引用 64 次
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 被引用 110 次
- Composable Augmentation Encoding for Video Representation LearningChen Sun, Arsha Nagrani, Yonglong Tian, Cordelia SchmidICCV 2021 · 被引用 20 次
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao 等ICCV 2021 · 被引用 37 次
