Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering
Sateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin, M. Zeeshan Zia, Quoc-Huy Tran
Abstract
We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequentially. We leverage temporal information in videos by employing temporal optimal transport. In particular, we incorporate a temporal regularization term which preserves the temporal order of the activity into the standard optimal transport module for computing pseudo-label cluster assignments. The temporal optimal transport module enables our approach to learn effective representations for unsupervised activity segmentation. Furthermore, previous methods require storing learned features for the entire dataset before clustering them in an offline manner, whereas our approach processes one mini-batch at a time in an online manner. Extensive evaluations on three public datasets, i.e. 50-Salads, YouTube Instructions, and Breakfast, and our dataset, i.e., Desktop Assembly, show that our approach performs on par with or better than previous methods, despite having significantly less memory constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9916e681-2bb7-4313-950b-e1a88edb9193Cited by top-tier papers18
- Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsZheng Xing, Weibing ZhaoAAAI 2024 · 18 citations
- Hierarchical Vector Quantization for Unsupervised Action SegmentationFederico Spurio, Emad Bahrami, Gianpiero Francesca, Juergen GallAAAI 2025 · 17 citations
- STREAMER: Streaming Representation Learning and Event Segmentation in a Hierarchical MannerRamy Mounir, Sujal Vijayaraghavan, Sudeep SarkarNeurIPS 2023 · 15 citations
- Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action SegmentationMing Xu, Stephen GouldCVPR 2024 · 15 citations
- OnlineTAS: An Online Baseline for Temporal Action SegmentationQing Zhong, Guodong Ding, Angela YaoNeurIPS 2024 · 15 citations
Builds on17
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
- Unsupervised Pre-Training of Image Features on Non-Curated DataMathilde Caron, Piotr Bojanowski, Julien Mairal, Armand JoulinICCV 2019 · 254 citations
Related papers
- CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationElena Belén Bueno-Benito, Mariella DimiccoliICCV 2025 · 3 citations
- Iterative Contrast-Classify for Semi-supervised Temporal Action SegmentationDipika Singhania, Rahul Rahaman, Angela YaoAAAI 2022 · 35 citations
- Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based ApproachQinying Liu, Zilei Wang, Shenghai Rong, Junjie Li et al.ICCV 2023 · 18 citations
- Action Shuffle Alternating Learning for Unsupervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2021
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin et al.ICCV 2025 · 13 citations
