TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition
Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah
摘要
Semi-Supervised Learning can be more beneficial for the video domain compared to images because of its higher annotation cost and dimensionality. Besides, any video understanding task requires reasoning over both spatial and temporal dimensions. In order to learn both the static and motion related features for the semi-supervised action recognition task, existing methods rely on hard input inductive biases like using two-modalities (RGB and Optical-flow) or two-stream of different playback rates. Instead of utilizing unlabeled videos through diverse input streams, we rely on self-supervised video representations, particularly, we utilize temporally-invariant and temporally-distinctive representations. We observe that these representations complement each other depending on the nature of the action. Based on this observation, we propose a student-teacher semi-supervised learning framework, TimeBalance, where we distill the knowledge from a temporally-invariant and a temporally-distinctive teacher. Depending on the nature of the unlabeled video, we dynamically combine the knowledge of these two teachers based on a novel temporal similarity-based reweighting scheme. Our method achieves state-of-the-art performance on three action recognition benchmarks: UCF101, HMDB51, and Kinetics400. Code: https://github . com/DAVEISHAN/TimeBalance. Classwise performance difference for action recognition task (% Accuracy of Temporally-Distinctive model) -(% Accuracy of Temporally-Invariant model) Temporally-Distinctive Representations perform better Clips share high-similarity, hence learning temporallyinvariant representation is useful Fencing Random Clip-1 Random Clip-2 Clips do not share high similarity, hence learning temporallydifferent(distinctive) representation is useful Random Clip-1 Random Clip-2 Javelin Throw Temporally-Invariant Representations perform better
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Semi-supervised Active Learning for Video Action DetectionAyush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas 等AAAI 2024 · 被引用 22 次
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 被引用 14 次
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 被引用 14 次
- SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationYongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia 等AAAI 2025 · 被引用 13 次
- MRI Contrast Enhancement Kinetics World ModelJindi Kong, Yuting He, Cong Xia, Rongjun Ge 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper40
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- Self-Supervised Learning for Semi-Supervised Temporal Action ProposalXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao 等CVPR 2021
- Learning from Temporal Gradient for Semi-supervised Action RecognitionJunfei Xiao, Longlong Jing, Lin Zhang, Ju He 等CVPR 2022 · 被引用 77 次
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 被引用 5 次
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah 等CVPR 2024 · 被引用 3 次
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 被引用 64 次
