Transcript to Video: Efficient Clip Sequencing from Texts
Yu Xiong, Fabian Caba Heilbron, Dahua Lin
摘要
Among numerous videos shared on the web, well-edited ones always attract more attention. However, it is difficult for inexperienced users to make well-edited videos because it requires professional expertise and immense manual labor. To meet the demands for non-experts, we present Transcript-to-Video -- a weakly-supervised framework that uses texts as input to automatically create video sequences from an extensive collection of shots. Specifically, we propose a Content Retrieval Module and a Temporal Coherent Module to learn visual-language representations and model shot sequencing styles, respectively. For fast inference, we introduce an efficient search strategy for real-time video clip sequencing. Quantitative results and user studies demonstrate empirically that the proposed learning framework can retrieve content-relevant shots while creating plausible video sequences in terms of style. Besides, the run-time performance analysis shows that our framework can support real-world applications. Project page: http://www.xiongyu.me/projects/transcript2video/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron 等CVPR 2022 · 被引用 84 次
- A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language ModelPanwen Hu, Nan Xiao, Feifei Li, Yongquan Chen 等ACM MM 2023 · 被引用 8 次
- EditDuet: A Multi-Agent System for Video Non-Linear EditingMarcelo Sandoval-Castañeda, Bryan C. Russell, Josef Sivic, Gregory Shakhnarovich 等SIGGRAPH 2025 · 被引用 6 次
- AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationMilton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen 等CVPR 2026 · 被引用 4 次
- Skald: Learning-Based Shot Assembly for Coherent Multi-Shot Video CreationChen-Yi Lu, Md. Mehrab Tanjim, Ishita Dasgupta, Somdeb Sarkhel 等ICCV 2025
它引用的顶会 Paper4
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 被引用 186 次
- End-to-End Learning of Visual Representations From Uncurated Instructional VideosAntoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev 等CVPR 2020
相关 Paper
- Self-supervised Video Summarization Guided by Semantic Inverse Optimal TransportYutong Wang, Hongteng Xu, Dixin LuoACM MM 2023 · 被引用 7 次
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion ModelsXiaoxue Wu, Bingjie Gao, Yu Qiao, Yaohui Wang 等ICLR 2026 · 被引用 26 次
- Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosSixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo 等CVPR 2023
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 被引用 1 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
