V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
Zhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding, Yali Wang
摘要
Despite the recent advancement in video stylization, most existing methods struggle to render any video with complex transitions, based on an open style description of user query. To fill this gap, we introduce a generic multi-agent system for video stylization, V-Stylist, by a novel collaboration and reflection paradigm of multi-modal large language models. Specifically, our V-Stylist is a systematical workflow with three key roles: (1) Video Parser decomposes the input video into a number of shots and generates their text prompts of key shot content. Via a concise video-to-shot prompting paradigm, it allows our V-Stylist to effectively handle videos with complex transitions. ( 2 ) Style Parser identifies the style in the user query and progressively search the matched style model from a style tree. Via a robust tree-of-thought searching paradigm, it allows our V-Stylist to precisely specify vague style preference in the open user query. (3) Style Artist leverages the matched model to render all the video shots into the required style. Via a novel multi-round self-reflection paradigm, it allows our V-Stylist to adaptively adjust detail control, according to the style requirement. With such a distinct design of mimicking human professionals, our V-Stylist achieves a major breakthrough over the primary challenges for effective and automatic video stylization. Moreover, we further construct a new benchmark Text-driven Video Stylization Benchmark (TVSBench), which fills the gap to assess various stylization of complex videos on open user queries. Extensive experiments show that, V-Stylist achieves the state-of-theart, e.g.,V-Stylist surpasses FRESCO and ControlVideo by 6.05% and 4.51% respectively in overall average metrics, marking a significant advance in video stylization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationZhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen 等ICLR 2026 · 被引用 25 次
- OASIS: On-Demand Hierarchical Event Memory for Streaming Video ReasoningZhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang 等CVPR 2026 · 被引用 16 次
- Video-GPT via Next Clip DiffusionShaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- GENMAC: Compositional Text-to-Video Generation with Multi-Agent CollaborationKaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin 等AAAI 2026 · 被引用 1 次
- Vlogger: Make Your Dream A VlogShaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang 等CVPR 2024 · 被引用 23 次
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
- Muses: 3D-Controllable Image Generation via Multi-Modal Agent CollaborationYanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue 等AAAI 2025 · 被引用 8 次
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang 等ICML 2026 · 被引用 8 次
