ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion Models
Hanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li, Chi-Sing Leung, Tien-Tsin Wong
Abstract
Video colorization is inherently challenging due to the need for accurate color inference and temporal consistency. In this paper, we present ColorDiffuser, an adaptation of a pre-trained text-to-image latent diffusion model for video colorization. By leveraging learned color priors from large-scale training, our method avoids costly retraining and enables controllable colorization via text prompts. To address the adaptation of an image model to video, we propose a novel Short- and Long-distance Cross-Frame Attention (SL-CFA) module combined with an amortized sampling strategy to unify the color latent over time. By incorporating information from nearby and distant frames, the model achieves better consistency for long video sequences, even with problematic disocclusion. To mitigate visual detail loss and color bleeding from compressed latent representations, we introduce a video colorization VAE model that incorporates semantic boundaries and grayscale inputs. Extensive experiments on benchmark datasets demonstrate that ColorDiffuser achieves state-of-the-art performance in color fidelity, temporal consistency, and visual quality, while offering diverse and controllable outputs. Our project page can be accessed at: https://colordiffuser.github.io.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d94e5fd5-686e-47f2-8d62-5da5257f1aa5Related papers
- Versatile Vision Foundation Model for Image and Video ColorizationVukasin Bozic, Abdelaziz Djelouah, Yang Zhang, Radu Timofte et al.SIGGRAPH 2024 · 9 citations
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsSijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang et al.NeurIPS 2024 · 82 citations
