Self-Supervised Motion Learning From Static Images
Ziyuan Huang, Shiwei Zhang, Jianwen Jiang, Mingqian Tang, Rong Jin, Marcelo H. Ang
Abstract
Motions are reflected in videos as the movement of pixels, and actions are essentially patterns of inconsistent motions between the foreground and the background. To well distinguish the actions, especially those with complicated spatio-temporal interactions, correctly locating the prominent motion areas is of crucial importance. However, most motion information in existing videos are difficult to label and training a model with good motion representations with supervision will thus require a large amount of human labour for annotation. In this paper, we address this problem by self-supervised learning. Specifically, we propose to learn Motion from Static Images (MoSI). The model learns to encode motion information by classifying pseudo motions generated by MoSI. We furthermore introduce a static mask in pseudo motions to create local motion patterns, which forces the model to additionally locate notable motion areas for the correct classification. We demonstrate that MoSI can discover regions with large motion even without finetuning on the downstream datasets. As a result, the learned motion representations boost the performance of tasks requiring understanding of complex scenes and motions, i.e., action recognition. Extensive experiments show the consistent and transferable improvements achieved by MoSI. Codes will be soon released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a662010-5fe9-43a7-a492-87f703dad864Cited by top-tier papers5
- TCTrack: Temporal Contexts for Aerial TrackingZiang Cao, Ziyuan Huang, Liang Pan, Shiwei Zhang et al.CVPR 2022 · 233 citations
- Tubelet-Contrastive Self-Supervision for Video-Efficient GeneralizationFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekICCV 2023 · 13 citations
- Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical ConsistencyZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu et al.CVPR 2022 · 11 citations
- SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningFida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard GhanemCVPR 2025
- Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in VideoRunyang Feng, Yixing Gao, Xueqing Ma, Tze Ho Elden Tse et al.CVPR 2023
Builds on9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
Related papers
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo et al.AAAI 2020 · 7 citations
- Learning Self-Similarity in Space and Time as Generalized Motion for Video Action RecognitionHeeseung Kwon, Manjin Kim, Suha Kwak, Minsu ChoICCV 2021 · 49 citations
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian et al.CVPR 2022 · 54 citations
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
