Self-supervised Video Summarization Guided by Semantic Inverse Optimal Transport
Yutong Wang, Hongteng Xu, Dixin Luo
Abstract
Video summarization is a critical task in video analysis that aims to create a brief yet informative summary of the original video (i.e., a set of keyframes) while retaining its primary content. Supervised summarization methods rely on time-consuming keyframe labeling and thus often suffer from the insufficiency issue of training data. In contrast, the performance of unsupervised summarization methods is often unsatisfactory due to the lack of semantically-meaningful guidance on the keyframe selection. In this study, we propose a novel self-supervised video summarization framework with the help of computational optimal transport techniques. Specifically, we generate textual descriptions from video shots and learn the projection from the textual embeddings to the visual ones together with an optimal transport plan between them via solving an inverse optimal transport problem. We propose an alternating optimization algorithm to solve this problem efficiently and design an effective mechanism in the algorithm to avoid trivial solutions. Given the optimal transport plan and the underlying distance between the projected textual embeddings and the visual ones, we synthesize pseudo-significance scores for video frames and leverage the scores as offline supervision to train a keyframe selector. Without subjective and error-prone manual annotations, the proposed framework surpasses previous unsupervised methods in producing high-quality results for generic and instructional video summarization tasks, whose performance even is comparable to those supervised competitors. The code is available at https://github.com/Dixin-s-Lab/Video-Summary-IOT.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25fa07f0-e5c0-4767-ae7d-a732ebc92330Cited by top-tier papers3
- VDOT: Efficient Unified Video Creation via Optimal Transport DistillationYutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao et al.CVPR 2026 · 5 citations
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer GenerationSidan Zhu, Hongteng Xu, Dixin LuoCVPR 2026 · 2 citations
- An Inverse Partial Optimal Transport Framework for Music-guided Trailer GenerationYutong Wang, Sidan Zhu, Hongteng Xu, Dixin LuoACM MM 2024 · 2 citations
Related papers
- Video Summarization Using Denoising Diffusion Probabilistic ModelZirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang et al.AAAI 2025 · 4 citations
- Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein DistanceDixin Luo, Yutong Wang, Angxiao Yue, Hongteng XuACM MM 2022 · 8 citations
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin et al.CVPR 2022 · 52 citations
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
