Preserving Global and Local Temporal Consistency for Arbitrary Video Style Transfer
Xinxiao Wu, Jialu Chen
Abstract
Video style transfer is a challenging task that requires not only stylizing video frames but also preserving temporal consistency among them. Many existing methods resort to optical flow for maintaining the temporal consistency in stylized videos. However, optical flow is sensitive to occlusions and rapid motions, and its training processing speed is quite slow, which makes it less practical in real-world applications. In this paper, we propose a novel fast method that explores both global and local temporal consistency for video style transfer without estimating optical flow. To preserve the temporal consistency of the entire video (i.e., global consistency), we use structural similarity index instead of flow optical and propose a self-similarity loss to ensure the temporal structure similarity between the stylized video and the source video. Furthermore, to enhance the coherence between adjacent frames (i.e., local consistency), a self-attention mechanism is designed to attend the previous stylized frame for synthesizing the current frame. Extensive experiments demonstrate that our method generally achieves better visual results and runs faster than the state-of-the-art methods, which validates the superiority of simultaneously preserving global and local temporal consistency for video style transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5b5e954-2365-435a-b132-0c527ca7bc8aCited by top-tier papers4
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li et al.ICCV 2021 · 421 citations
- Two Birds, One Stone: A Unified Framework for Joint Learning of Image and Video Style TransfersBohai Gu, Heng Fan, Libo ZhangICCV 2023 · 20 citations
- Video Color Grading via Look-Up Table GenerationSeunghyun Shin, Dongmin Shin, Jisu Shin, Hae-Gon Jeon et al.ICCV 2025 · 2 citations
- ReGS: Reference-based Controllable Scene Stylization with Gaussian SplattingYiqun Mei, Jiacong Xu, Vishal M. PatelNeurIPS 2024
Builds on1
Related papers
- Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyZhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin et al.AAAI 2022 · 7 citations
- Stable Video Style Transfer Based on Partial Convolution with Depth-Aware SupervisionSonghua Liu, Hao Wu, Shoutong Luo, Zhengxing SunACM MM 2020 · 6 citations
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi et al.AAAI 2020 · 80 citations
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 7 citations
- Time Flies: Animating a Still Image With Time-Lapse Video As ReferenceChia-Chi Cheng, Hung-Yu Chen, Wei-Chen ChiuCVPR 2020
