Video Summarization Using Denoising Diffusion Probabilistic Model
Zirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang, Xinxiao Wu
Abstract
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative models to predict the importance scores of video frames. However, these methods are susceptible to annotation inconsistency caused by the inherent subjectivity of different annotators when annotating the same video. In this paper, we introduce a generative framework for video summarization that learns how to generate summaries from a probability distribution perspective, effectively reducing the interference of subjective annotation noise. Specifically, we propose a novel diffusion summarization method based on the Denoising Diffusion Probabilistic Model (DDPM), which learns the probability distribution of training data through noise prediction, and generates summaries by iterative denoising. Our method is more resistant to subjective annotation noise, and is less prone to overfitting the training data than discriminative methods, with strong generalization ability. Moreover, to facilitate training DDPM with limited data, we employ an unsupervised video summarization model to implement the earlier denoising process. Extensive experiments on various datasets (TVSum, SumMe, and FPVSum) demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baf439c4-9dc5-4c9e-b709-61d6b4c34b87Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- A Latent Space of Stochastic Diffusion Models for Zero-Shot Image Editing and GuidanceChen Henry Wu, Fernando De la TorreICCV 2023 · 141 citations
- Implicit Diffusion Models for Continuous Super-ResolutionSicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu et al.CVPR 2023
Related papers
- SummDiff: Generative Modeling of Video Summarization with DiffusionKwanseok Kim, Jaehoon Hahm, Sumin Kim, Jinhwan Sul et al.ICCV 2025 · 1 citation
- VideoFusion: Decomposed Diffusion Models for High-Quality Video GenerationCVPR 2023
- Feature Prediction Diffusion Model for Video Anomaly DetectionCheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang et al.ICCV 2023 · 76 citations
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationShuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu et al.CVPR 2025
- Self-supervised Video Summarization Guided by Semantic Inverse Optimal TransportYutong Wang, Hongteng Xu, Dixin LuoACM MM 2023 · 7 citations
