Time Is MattEr: Temporal Self-supervision for Video Transformers
Sukmin Yun, Jaehyung Kim, Dongyoon Han, Hwanjun Song, Jung-Woo Ha, Jinwoo Shin
Abstract
Understanding temporal dynamics of video is an essential aspect of learning better video representations. Recently, transformer-based architectural designs have been extensively explored for video tasks due to their capability to capture long-term dependency of input sequences. However, we found that these Video Transformers are still biased to learn spatial dynamics rather than temporal ones, and debiasing the spurious correlation is critical for their performance. Based on the observations, we design simple yet effective self-supervised tasks for video models to learn temporal dynamics better. Specifically, for debiasing the spatial bias, our method learns the temporal order of video frames as extra self-supervision and enforces the randomly shuffled frames to have low-confidence outputs. Also, our method learns the temporal flow direction of video tokens among consecutive frames for enhancing the correlation toward temporal dynamics. Under various video action recognition tasks, we demonstrate the effectiveness of our method and its compatibility with state-of-the-art Video Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Mitigating and Evaluating Static Bias of Action Representations in the Background and the ForegroundHaoxin Li, Yuan Liu, Hanwang Zhang, Boyang LiICCV 2023 · 30 citations
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui et al.ACM MM 2023 · 15 citations
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 14 citations
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya et al.ICLR 2025 · 2 citations
- Soft-Landing Strategy for Alleviating the Task Discrepancy Problem in Temporal Action Localization TasksHyolim Kang, Hanjung Kim, Joungbin An, Minsu Cho et al.CVPR 2023
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan et al.CVPR 2022 · 111 citations
- Cross-Architecture Self-supervised Video Representation LearningSheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang et al.CVPR 2022 · 23 citations
- Video Language Model Pretraining with Spatio-temporal MaskingYue Wu, Zhaobo Qi, Junshu Sun, Yaowei Wang et al.CVPR 2025
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj et al.ICCV 2021 · 76 citations
