Self-supervised Video Summarization Guided by Semantic Inverse Optimal Transport
Yutong Wang, Hongteng Xu, Dixin Luo
摘要
Video summarization is a critical task in video analysis that aims to create a brief yet informative summary of the original video (i.e., a set of keyframes) while retaining its primary content. Supervised summarization methods rely on time-consuming keyframe labeling and thus often suffer from the insufficiency issue of training data. In contrast, the performance of unsupervised summarization methods is often unsatisfactory due to the lack of semantically-meaningful guidance on the keyframe selection. In this study, we propose a novel self-supervised video summarization framework with the help of computational optimal transport techniques. Specifically, we generate textual descriptions from video shots and learn the projection from the textual embeddings to the visual ones together with an optimal transport plan between them via solving an inverse optimal transport problem. We propose an alternating optimization algorithm to solve this problem efficiently and design an effective mechanism in the algorithm to avoid trivial solutions. Given the optimal transport plan and the underlying distance between the projected textual embeddings and the visual ones, we synthesize pseudo-significance scores for video frames and leverage the scores as offline supervision to train a keyframe selector. Without subjective and error-prone manual annotations, the proposed framework surpasses previous unsupervised methods in producing high-quality results for generic and instructional video summarization tasks, whose performance even is comparable to those supervised competitors. The code is available at https://github.com/Dixin-s-Lab/Video-Summary-IOT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- VDOT: Efficient Unified Video Creation via Optimal Transport DistillationYutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao 等CVPR 2026 · 被引用 5 次
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer GenerationSidan Zhu, Hongteng Xu, Dixin LuoCVPR 2026 · 被引用 2 次
- An Inverse Partial Optimal Transport Framework for Music-guided Trailer GenerationYutong Wang, Sidan Zhu, Hongteng Xu, Dixin LuoACM MM 2024 · 被引用 2 次
相关 Paper
- Video Summarization Using Denoising Diffusion Probabilistic ModelZirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang 等AAAI 2025 · 被引用 4 次
- Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein DistanceDixin Luo, Yutong Wang, Angxiao Yue, Hongteng XuACM MM 2022 · 被引用 8 次
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin 等CVPR 2022 · 被引用 52 次
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 被引用 196 次
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
