Modeling Video as Stochastic Processes for Fine-Grained Video Representation Learning
Heng Zhang, Daqing Liu, Qi Zheng, Bing Su
Abstract
A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video. In this paper, we propose to learn video representations by modeling Video as Stochastic Processes (VSP) via a novel process-based contrastive learning framework, which aims to discriminate between video processes and simultaneously capture the temporal dynamics in the processes. Specifically, we enforce the embeddings of the frame sequence of interest to approximate a goal-oriented stochastic process, i.e., Brownian bridge, in the latent space via a process-based contrastive loss. To construct the Brownian bridge, we adapt specialized sampling strategies under different annotations for both self-supervised and weakly-supervised learning. Experimental results on four datasets show that VSP stands as a state-of-the-art method for various video understanding tasks, including phase progression, phase classification and frame retrieval. Code is available at 'https://github.com/hengRUC/VSP'.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang et al.CVPR 2024 · 31 citations
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 14 citations
- SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMsChong Tang, Sannara Ek, Dirk Koch, Robert Mullins et al.ICLR 2026
- BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic RepresentationTianhao Zhang, Zhecheng Sheng, Zhexiao Lin, Chen Jiang et al.EMNLP 2025
- Synchronization of Multiple VideosAvihai Naaman, Ron Shapira Weber, Oren FreifeldICCV 2025
Builds on13
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long et al.AAAI 2021 · 140 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 34 citations
Related papers
- Probabilistic Representations for Video Contrastive LearningJungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon SohnCVPR 2022 · 42 citations
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 6 citations
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
- Composable Augmentation Encoding for Video Representation LearningChen Sun, Arsha Nagrani, Yonglong Tian, Cordelia SchmidICCV 2021 · 20 citations
- Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation LearningYuan Yao, Chang Liu, Dezhao Luo, Yu Zhou et al.CVPR 2020
