Vlogger: Make Your Dream A Vlog
Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang, Ziwei Liu, Yu Qiao, Yali Wang
摘要
In this work, we present Vlogger, a generic AI systemfor generating a minute-level video blog (i.e., vlog) of user de-scriptions. Different from short videos with a few seconds, vlog often contains a complex storyline with diversified scenes, which is challenging for most existing video generation approaches. To break through this bottleneck, our Vlogger smartly leverages Large Language Model (LLM) as Director and decomposes a long video generation task of vlog into four key stages, where we invoke various foundation models to play the critical roles of vlog profession-als, including (1) Script, (2) Actor, (3) ShowMaker, and (4) Voicer. With such a design of mimicking human beings, our Vlogger can generate vlogs through explainable cooperation of top-down planning and bottom-up shooting. More-over, we introduce a novel video diffusion model, Show-Maker, which serves as a videographer in our Vlogger for generating the video snippet of each shooting scene. By incorporating Script and Actor attentively as textual and visual prompts, it can effectively enhance spatial-temporal coherence in the snippet. Besides, we design a concise mixed training paradigm for ShowMaker, boosting its ca-pacity for both T2V generation and prediction. Finally, the extensive experiments show that our method achieves state-of-the-art performance on zero-shot T2V generation and prediction tasks. More importantly, Vlogger can generate over 5-minute vlogs from open-world descriptions, without loss of video coherence on script and actor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao 等NeurIPS 2024 · 被引用 62 次
- Captain Cinema: Towards Short Movie GenerationJunfei Xiao, Ceyuan Yang, Lvmin Zhang, Shengqu Cai 等ICLR 2026 · 被引用 51 次
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu 等CVPR 2026 · 被引用 37 次
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang 等ICLR 2026 · 被引用 26 次
- EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationRang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng 等AAAI 2026 · 被引用 24 次
它引用的顶会 Paper34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- V-Stylist: Video Stylization via Collaboration and Reflection of MLLM AgentsZhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding 等CVPR 2025
- Bridging Your Imagination with Audio-Video Generation via a Unified DirectorJiaxu Zhang, Tianshu Hu, Yuan Zhang, Zenan Li 等ICML 2026 · 被引用 2 次
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- LLM-grounded Video Diffusion ModelsLong Lian, Baifeng Shi, Adam Yala, Trevor Darrell 等ICLR 2024 · 被引用 87 次
- SEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionXinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang 等ICLR 2024 · 被引用 226 次
