Towards Automated Movie Trailer Generation
Dawit Mureja Argaw, Mattia Soldan, Alejandro Pardo, Chen Zhao, Fabian Caba Heilbron, Joon Son Chung, Bernard Ghanem
Abstract
Movie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To streamline this process, we propose an automatic trailer generation framework that generates plausible trailers from a full movie by automating shot selection and composition. Our approach draws inspiration from machine translation techniques and models the movies and trailers as sequences of shots, thus formulating the trailer generation problem as a sequence-to-sequence task. We introduce Trailer Generation Transformer (TGT), a deep-learning framework utilizing an encoder-decoder architecture. TGT movie encoder is tasked with contextualizing each movie shot representation via self-attention, while the autoregressive trailer decoder predicts the feature representation of the next trailer shot, accounting for the relevance of shots' temporal order in trailers. Our TGT significantly outperforms previous methods on a comprehensive suite of metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c78951c4-07d5-408d-9fcf-511a74b2930dCited by top-tier papers4
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer GenerationSidan Zhu, Hongteng Xu, Dixin LuoCVPR 2026 · 2 citations
- REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video EditingWeihan Xu, Yimeng Ma, Jingyue Huang, Yang Li et al.NeurIPS 2025 · 1 citation
- Autoregressive Modeling of Film with Applications in Video MontageMarcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar, Bryan C. Russell et al.SIGGRAPH 2026
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- Contrastive Learning for Unsupervised Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengCVPR 2022 · 39 citations
- Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in MoviesBei Gan, Xiujun Shu, Ruizhi Qiao, Haoqian Wu et al.CVPR 2023
Related papers
- An Inverse Partial Optimal Transport Framework for Music-guided Trailer GenerationYutong Wang, Sidan Zhu, Hongteng Xu, Dixin LuoACM MM 2024 · 2 citations
- Teaching Temporal Logics to Neural NetworksChristopher Hahn, Frederik Schmitt, Jens U. Kreber, Markus Norman Rabe et al.ICLR 2021 · 78 citations
- Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description GenerationJunyu Xie, Tengda Han, Max Bain, Arsha Nagrani et al.ICCV 2025
- Narrative Plan Generation with Self-Supervised LearningMihai Polceanu, Julie Porteous, Alan Lindsay, Marc CavazzaAAAI 2021 · 5 citations
- Long Context Tuning for Video GenerationYuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma et al.ICCV 2025 · 6 citations
