FreeViS: Training-free Video Stylization with Inconsistent References
Jiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. Patel
Abstract
Video stylization plays a key role in content creation, but it remains a challenging problem. Naïvely applying image stylization frame-by-frame hurts temporal consistency and reduces style richness. Alternatively, training a dedicated video stylization model typically requires paired video data and is computationally expensive. In this paper, we propose FreeViS, a training-free video stylization framework that generates stylized videos with rich style details and strong temporal coherence. Our method integrates multiple stylized references to a pretrained image-to-video (I2V) model, effectively mitigating the propagation errors observed in prior works, without introducing flickers and stutters. In addition, it leverages high-frequency compensation to constrain the content layout and motion, together with flow-based motion cues to preserve style textures in low-saliency regions. Through extensive evaluations, FreeViS delivers higher stylization fidelity and superior temporal consistency, outperforming recent baselines and achieving strong human preference. Our training-free pipeline offers a practical and economic solution for high-quality, temporally coherent video stylization. The code and videos can be accessed via https://xujiacong.github.io/FreeViS/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7a408bf-a767-4f22-9090-9bcd0b0c9315Cited by top-tier papers1
Ask how each one uses itBuilds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- Preserving Global and Local Temporal Consistency for Arbitrary Video Style TransferXinxiao Wu, Jialu ChenACM MM 2020 · 14 citations
- VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion ModelsYabo Zhang, Yuxiang Wei, Xianhui Lin, Zheng Hui et al.AAAI 2025 · 3 citations
- Interactive video stylization using few-shot patch-based trainingOndrej Texler, David Futschik, Michal Kucera, Ondrej Jamriska et al.SIGGRAPH 2020 · 70 citations
- Stable Video Style Transfer Based on Partial Convolution with Depth-Aware SupervisionSonghua Liu, Hao Wu, Shoutong Luo, Zhengxing SunACM MM 2020 · 6 citations
- FlowMotion: Training-Free Flow Guidance for Video Motion TransferZhen Wang, Youcan Xu, Jun Xiao, Long ChenCVPR 2026 · 1 citation
