An Inverse Partial Optimal Transport Framework for Music-guided Trailer Generation
Yutong Wang, Sidan Zhu, Hongteng Xu, Dixin Luo
Abstract
Trailer generation is a challenging video clipping task that aims to select highlighting shots from long videos like movies and re-organize them in an attractive way. In this study, we propose an inverse partial optimal transport (IPOT) framework to achieve music-guided movie trailer generation. In particular, we formulate the trailer generation task as selecting and sorting key movie shots based on audio shots, which involves matching the latent representations across visual and acoustic modalities. We learn a multi-modal latent representation model in the proposed IPOT framework to achieve this aim. In this framework, a two-tower encoder derives the latent representations of movie and music shots, respectively, and an attention-assisted Sinkhorn matching network parameterizes the grounding distance between the shots' latent representations and the distribution of the movie shots. Taking the correspondence between the movie shots and its trailer music shots as the observed optimal transport plan defined on the grounding distances, we learn the model by solving an inverse partial optimal transport problem, leading to a bi-level optimization strategy. We collect real-world movies and their trailers to construct a dataset with abundant label information called CMTD and, accordingly, train and evaluate various automatic trailer generators. Compared with state-of-the-art methods, our IPOT method consistently shows superiority in subjective visual effects and objective quantitative measurements. The code is available at https://github.com/Dixin-Lab/Automatic-Movie-Trailer-Generator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d39c40b1-a50c-4a06-81ba-bfa2da647b6cCited by top-tier papers2
- VDOT: Efficient Unified Video Creation via Optimal Transport DistillationYutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao et al.CVPR 2026 · 5 citations
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer GenerationSidan Zhu, Hongteng Xu, Dixin LuoCVPR 2026 · 2 citations
Builds on16
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 · 193 citations
- Accurate Point Cloud Registration with Robust Optimal TransportZhengyang Shen, Jean Feydy, Peirong Liu, Ariel Hernán Curiale et al.NeurIPS 2021 · 81 citations
- Gromov-Wasserstein Factorization Models for Graph ClusteringHongteng XuAAAI 2020 · 56 citations
Related papers
- Towards Automated Movie Trailer GenerationDawit Mureja Argaw, Mattia Soldan, Alejandro Pardo, Chen Zhao et al.CVPR 2024
- Self-supervised Video Summarization Guided by Semantic Inverse Optimal TransportYutong Wang, Hongteng Xu, Dixin LuoACM MM 2023 · 7 citations
- VMChill: A Dataset for Fine-Grained Visual-Musical SynergyXiaowei Chi, Zeyue Tian, Jialiang Chen, Wei XueAAAI 2026
- TeaserGen: Generating Teasers for Long DocumentariesWeihan Xu, Paul Pu Liang, Haven Kim, Julian J. McAuley et al.ICLR 2025
- OT-CLIP: Understanding and Generalizing CLIP via Optimal TransportLiangliang Shi, Jack Fan, Junchi YanICML 2024 · 11 citations
