Learning to Cut by Watching Movies
Alejandro Pardo, Fabian Caba Heilbron, Juan León Alcázar, Ali K. Thabet, Bernard Ghanem
Abstract
Video content creation keeps growing at an incredible pace; yet, creating engaging stories remains challenging and requires non-trivial video editing expertise. Many video editing components are astonishingly hard to automate primarily due to the lack of raw video materials. This paper focuses on a new task for computational video editing, namely the task of raking cut plausibility. Our key idea is to leverage content that has already been edited to learn fine-grained audiovisual patterns that trigger cuts. To do this, we first collected a data source of more than 10K videos, from which we extract more than 255K cuts. We devise a model that learns to discriminate between real and artificial cuts via contrastive learning. We set up a new task and a set of baselines to benchmark video cut generation. We observe that our proposed model outperforms the baselines by large margins. To demonstrate our model in real-world applications, we conduct human studies in a collection of unedited videos. The results show that our model does a better job at cutting than random and alternative baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d155b3e9-cec5-4dcc-abd0-d81becdfe945Cited by top-tier papers5
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- Long-range Multimodal Pretraining for Movie UnderstandingDawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon et al.ICCV 2023 · 15 citations
- ChunkyEdit: Text-first video interview editing via chunkingMackenzie Leake, Wilmot LiCHI 2024 · 8 citations
- MatchDiffusion: Training-Free Generation of Match-CutsAlejandro Pardo, Fabio Pizzati, Tong Zhang, Alexander Pondaven et al.ICCV 2025 · 3 citations
- VEU-Bench: Towards Comprehensive Understanding of Video EditingBozheng Li, Yongliang Wu, Yi Lu, Jiashuo Yu et al.CVPR 2025
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
Related papers
- Contrastive Learning for Unsupervised Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengCVPR 2022 · 39 citations
- VideoCon: Robust Video-Language Alignment via Contrast CaptionsHritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang et al.CVPR 2024
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi et al.NeurIPS 2023 · 26 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley et al.ICLR 2023 · 3 citations
