Transcript to Video: Efficient Clip Sequencing from Texts
Yu Xiong, Fabian Caba Heilbron, Dahua Lin
Abstract
Among numerous videos shared on the web, well-edited ones always attract more attention. However, it is difficult for inexperienced users to make well-edited videos because it requires professional expertise and immense manual labor. To meet the demands for non-experts, we present Transcript-to-Video -- a weakly-supervised framework that uses texts as input to automatically create video sequences from an extensive collection of shots. Specifically, we propose a Content Retrieval Module and a Temporal Coherent Module to learn visual-language representations and model shot sequencing styles, respectively. For fast inference, we introduce an efficient search strategy for real-time video clip sequencing. Quantitative results and user studies demonstrate empirically that the proposed learning framework can retrieve content-relevant shots while creating plausible video sequences in terms of style. Besides, the run-time performance analysis shows that our framework can support real-world applications. Project page: http://www.xiongyu.me/projects/transcript2video/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c29ee9a-37da-4474-9107-d80cf2336951Cited by top-tier papers6
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language ModelPanwen Hu, Nan Xiao, Feifei Li, Yongquan Chen et al.ACM MM 2023 · 8 citations
- EditDuet: A Multi-Agent System for Video Non-Linear EditingMarcelo Sandoval-Castañeda, Bryan C. Russell, Josef Sivic, Gregory Shakhnarovich et al.SIGGRAPH 2025 · 6 citations
- AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationMilton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen et al.CVPR 2026 · 4 citations
- Skald: Learning-Based Shot Assembly for Coherent Multi-Shot Video CreationChen-Yi Lu, Md. Mehrab Tanjim, Ishita Dasgupta, Somdeb Sarkhel et al.ICCV 2025
Builds on4
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 186 citations
- End-to-End Learning of Visual Representations From Uncurated Instructional VideosAntoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev et al.CVPR 2020
Related papers
- Self-supervised Video Summarization Guided by Semantic Inverse Optimal TransportYutong Wang, Hongteng Xu, Dixin LuoACM MM 2023 · 7 citations
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion ModelsXiaoxue Wu, Bingjie Gao, Yu Qiao, Yaohui Wang et al.ICLR 2026 · 26 citations
- Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosSixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo et al.CVPR 2023
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 1 citation
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel et al.ICCV 2023 · 800 citations
