SLIC: Self-Supervised Learning with Iterative Clustering for Human Action Videos
Salar Hosseini Khorasgani, Yuxuan Chen, Florian Shkurti
Abstract
Self-supervised methods have significantly closed the gap with end-to-end supervised learning for image classification [13], [24]. In the case of human action videos, however, where both appearance and motion are significant factors of variation, this gap remains significant [28], [58]. One of the key reasons for this is that sampling pairs of similar video clips, a required step for many self-supervised contrastive learning methods, is currently done conservatively to avoid false positives. A typical assumption is that similar clips only occur temporally close within a single video, leading to insufficient examples of motion similarity. To mitigate this, we propose SLIC, a clustering-based self-supervised contrastive learning method for human action videos. Our key contribution is that we improve upon the traditional intra-video positive sampling by using iterative clustering to group similar video instances. This enables our method to leverage pseudo-labels from the cluster assignments to sample harder positives and negatives. SLIC outperforms state-of-the-art video retrieval baselines by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> on top-1 recall on UCF101 and by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> when directly transferred to HMDB51. With end-to-end finetuning for action classi-fication, SLIC achieves 83.2% top-1 accuracy <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> on UCF101 and 54.5% on HMDB51 <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> ,. SLIC is also competitive with the state-of-the-art in action classification after self-supervised pretraining on Kinetics400.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time SeriesChenxi Sun, Hongyan Li, Yaliang Li, Shenda HongICLR 2024 · 223 citations
- MHCCL: Masked Hierarchical Cluster-Wise Contrastive Learning for Multivariate Time SeriesQianwen Meng, Hangwei Qian, Yong Liu, Lizhen Cui et al.AAAI 2023 · 54 citations
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 14 citations
- Cross-modal Scalable Hyperbolic Hierarchical ClusteringTeng Long, Nanne van NoordICCV 2023 · 12 citations
- Region-Based Representations RevisitedMichal Shlapentokh-Rothman, Ansel Blume, Yao Xiao, Yuqun Wu et al.CVPR 2024 · 11 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
- Probabilistic Representations for Video Contrastive LearningJungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon SohnCVPR 2022 · 42 citations
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu et al.ICCV 2021 · 55 citations
- Self-Supervised Video Representation Learning with Meta-Contrastive NetworkYuanze Lin, Xun Guo, Yan LuICCV 2021 · 46 citations
