Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein Distance
Dixin Luo, Yutong Wang, Angxiao Yue, Hongteng Xu
Abstract
Temporal action alignment aims at segmenting videos into clips and tagging each clip with a textual description, which is an important task of video semantic analysis. Most existing methods, however, rely on supervised learning to train their alignment models, whose applications are limited because of the common insufficiency issue of labeled videos. To mitigate this issue, we propose a weakly-supervised temporal action alignment method based on a novel computational optimal transport technique called unbalanced spectral fused Gromov-Wasserstein (US-FGW) distance. Instead of using videos with known clips and corresponding textual tags, our method just needs each training video to be associated with a set of (unsorted) texts while does not require the fine-grained correspondence between the frames and the texts. Given such weakly-supervised video-text pairs, our method trains the representation models of the video frames and the texts jointly in a probabilistic or deterministic autoencoding architecture and penalizes the US-FGW distance between the distribution of visual latent codes and that of textual latent codes. We compute the US-FGW distance efficiently by leveraging the Bregman ADMM algorithm. Furthermore, we generalize classic contrastive learning framework and reformulate it based on the proposed US-FGW distance, which provides a new viewpoint of contrastive learning for our problem. Experimental results show that our method and its variants outperform state-of-the-art weakly-supervised temporal action alignment methods, whose results are even comparable to those derived by supervised learning methods on some specific evaluation measurements. The code is available at ://github.com/hhhh1138/Temporal-Action-Alignment-USFGW.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 04ea13ef-a067-4295-9e40-636cf0c303eeCited by top-tier papers3
- Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsZheng Xing, Weibing ZhaoAAAI 2024 · 18 citations
- VDOT: Efficient Unified Video Creation via Optimal Transport DistillationYutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao et al.CVPR 2026 · 5 citations
- An Inverse Partial Optimal Transport Framework for Music-guided Trailer GenerationYutong Wang, Sidan Zhu, Hongteng Xu, Dixin LuoACM MM 2024 · 2 citations
Related papers
- Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action SegmentationMing Xu, Stephen GouldCVPR 2024 · 15 citations
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin et al.ICCV 2025 · 13 citations
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh et al.CVPR 2022 · 18 citations
- Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based ApproachQinying Liu, Zilei Wang, Shenghai Rong, Junjie Li et al.ICCV 2023 · 18 citations
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 11 citations
