Autoregressive Modeling of Film with Applications in Video Montage
Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar, Bryan C. Russell, Josef Sivic, Alexei A. Efros, Gregory Shakhnarovich
Abstract
This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage – turning a collection of raw, “unwatchable” footage into coherent cinematic sequences. Inspired by language learning in modern LLMs, we train a long-context autoregressive transformer on a large corpus of movies. The aim is to implicitly capture the “grammar” of film directly from data rather than from hand-coded rules. Unlike other generative models, FilmGPT does not generate any new video frames. Instead, at inference time, we introduce a footage-constrained decoding algorithm to select the best next shot from the input raw footage according to the statistical patterns learned from films. We first evaluate these learned statistics directly by using the FilmGPT autoregressive model for next shot prediction on a standard benchmark of shot sequence ordering, outperforming the previous state of the art. We then evaluate our footage-constrained decoding algorithm on the full film editing task via a user study, and find that our FilmGPT-based editing significantly outperforms previous approaches. Finally, we demonstrate the applicability of FilmGPT to a wide range of applications in video montage, from automatic video segment trimming to human-in-the-loop film editing. Please see our supplementary material for qualitative results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e1f2c70-fe0d-4bff-9772-a8ec46396d41Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen et al.CHI 2023 · 44 citations
- GAZED- Gaze-guided Cinematic Editing of Wide-Angle Monocular Video RecordingsK. L. Bhanu Moorthy, Moneish Kumar, Ramanathan Subramanian, Vineet GandhiCHI 2020 · 26 citations
Related papers
- Towards Automated Movie Trailer GenerationDawit Mureja Argaw, Mattia Soldan, Alejandro Pardo, Chen Zhao et al.CVPR 2024
- Video-GPT via Next Clip DiffusionShaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang et al.ICLR 2026 · 9 citations
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li et al.ICCV 2025 · 5 citations
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer GenerationSidan Zhu, Hongteng Xu, Dixin LuoCVPR 2026 · 2 citations
- End-to-end Generative Pretraining for Multimodal Video CaptioningPaul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia SchmidCVPR 2022 · 152 citations
