Skald: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
Chen-Yi Lu, Md. Mehrab Tanjim, Ishita Dasgupta, Somdeb Sarkhel, Gang Wu, Saayan Mitra, Somali Chaterji
Abstract
We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Learned Clip Assembly (LCA) score, a learning-based metric that measures temporal and semantic relationships between shots to quantify narrative coherence. We tackle the exponential complexity of combining multiple shots with an efficient beam-search algorithm guided by the LCA score. To train our model effectively with limited human annotations, we propose two tasks for the LCA encoder: Shot Coherence Learning, which uses contrastive learning to distinguish coherent and incoherent sequences, and Feature Regression, which converts these learned representations into a real-valued coherence score. We develop two variants: a base SKALD model that relies solely on visual coherence and SKALD-text, which integrates auxiliary text information when available. Experiments on the VSPD and our curated MSV3C datasets show that SKALD achieves an improvement of up to 48.6% in IoU and a 43% speedup over the state-of-the-art methods. A user study further validates our approach, with 45% of participants favoring SKALD-assembled videos, compared to 22% preferring text-based assembly methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b3ea54f-3207-4fca-90f6-e42c61c78e5bCited by top-tier papers1
Ask how each one uses itBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen et al.ICML 2024 · 499 citations
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ICCV 2023 · 371 citations
Related papers
- Positive-Augmented Contrastive Learning for Image and Video Captioning EvaluationSara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2023
- TACo: Token-aware Cascade Contrastive Learning for Video-Text AlignmentJianwei Yang, Yonatan Bisk, Jianfeng GaoICCV 2021 · 159 citations
- RCA-NOC: Relative Contrastive Alignment for Novel Object CaptioningJiashuo Fan, Yaoyuan Liang, Leyao Liu, Shao-Lun Huang et al.ICCV 2023 · 7 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description GenerationJunyu Xie, Tengda Han, Max Bain, Arsha Nagrani et al.ICCV 2025
