Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action Segmentation
Ming Xu, Stephen Gould
Abstract
We propose a novel approach to the action segmentation task for long, untrimmed videos, based on solving an optimal transport problem. By encoding a temporal consistency prior into a Gromov-Wasserstein problem, we are able to decode a temporally consistent segmentation from a noisy affinity/matching cost matrix between video frames and action classes. Unlike previous approaches, our method does not require knowing the action order for a video to attain temporal consistency. Furthermore, our resulting (fused) Gromov-Wasserstein problem can be efficiently solved on GPUs using a few iterations of projected mirror descent. We demonstrate the effectiveness of our method in an unsupervised learning setting, where our method is used to generate pseudo-labels for self-training. We evaluate our segmentation approach and unsupervised learning pipeline on the Breakfast, 50-Salads, YouTube Instructions and Desktop Assembly datasets, yielding state-of-the-art results for the unsupervised video action segmentation task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c73a35f-1b14-4ae4-8da7-197c33209778Cited by top-tier papers13
- Hierarchical Vector Quantization for Unsupervised Action SegmentationFederico Spurio, Emad Bahrami, Gianpiero Francesca, Juergen GallAAAI 2025 · 17 citations
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin et al.ICCV 2025 · 13 citations
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
- CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationElena Belén Bueno-Benito, Mariella DimiccoliICCV 2025 · 3 citations
- Motion Control via Metric-Aligning Motion MatchingNaoki Agata, Takeo IgarashiSIGGRAPH 2025 · 2 citations
Builds on15
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- The Unbalanced Gromov Wasserstein Distance: Conic Formulation and RelaxationThibault Séjourné, François-Xavier Vialard, Gabriel PeyréNeurIPS 2021 · 106 citations
- Accurate Point Cloud Registration with Robust Optimal TransportZhengyang Shen, Jean Feydy, Peirong Liu, Ariel Hernán Curiale et al.NeurIPS 2021 · 81 citations
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 74 citations
Related papers
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin et al.CVPR 2022 · 52 citations
- Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein DistanceDixin Luo, Yutong Wang, Angxiao Yue, Hongteng XuACM MM 2022 · 8 citations
- Action Shuffle Alternating Learning for Unsupervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2021
- Learning to Align Sequential Actions in the WildWeizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet et al.CVPR 2022
- Iterative Contrast-Classify for Semi-supervised Temporal Action SegmentationDipika Singhania, Rahul Rahaman, Angela YaoAAAI 2022 · 35 citations
