Frame-wise Action Representations for Long Videos via Sequence Contrastive Learning
Minghao Chen, Fangyun Wei, Chong Li, Deng Cai
摘要
Prior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practical applications such as video alignment have strong demand for learning dense representations for long videos. In this paper, we introduce a novel contrastive action representation learning (CARL) framework to learn frame-wise action representations, especially for long videos, in a self-supervised manner. Concretely, we introduce a simple yet efficient video encoder that considers spatio-temporal context to extract frame-wise representations. Inspired by the recent progress of self-supervised learning, we present a novel sequence contrastive loss (SCL) applied on two correlated views obtained through a series of spatio-temporal data augmentations. SCL optimizes the embedding space by minimizing the KL-divergence between the sequence similarity of two augmented views and a prior Gaussian distribution of timestamp distance. Experiments on FineGym, PennAction and Pouring datasets show that our method outperforms previous state-of-the-art by a large margin for downstream fine-grained action classification. Surprisingly, although without training on paired videos, our approach also shows outstanding performance on video alignment and fine-grained frame retrieval tasks. Code and models are available at https://github.com/minghchen/CARL_code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li 等ICCV 2023 · 被引用 77 次
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 被引用 64 次
- Conditional Information Bottleneck Approach for Time Series ImputationMinGyu Choi, Changhee LeeICLR 2024 · 被引用 24 次
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 18 次
- ViSTec: Video Modeling for Sports Technique Recognition and Tactical AnalysisYuchen He, Zeqing Yuan, Yihong Wu, Liqi Cheng 等AAAI 2024 · 被引用 13 次
它引用的顶会 Paper18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation LearningBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ACM MM 2022 · 被引用 6 次
- Composable Augmentation Encoding for Video Representation LearningChen Sun, Arsha Nagrani, Yonglong Tian, Cordelia SchmidICCV 2021 · 被引用 20 次
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao 等ACM MM 2021 · 被引用 6 次
- Cross-Architecture Self-supervised Video Representation LearningSheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang 等CVPR 2022 · 被引用 23 次
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang 等CVPR 2021
