Towards Photorealistic Video Colorization via Gated Color-Guided Image Diffusion Models
Jiaxing Li, Hongbo Zhao, Yijun Wang, Jianxin Lin
摘要
Video colorization poses challenging tasks, necessitating structural stability, continuity, and details control in the colors produced. In this paper, based on a pretrained text-to-image model, we introduce the Gated Color Guidance module (GCG ), enabling the model to adaptively perform color propagation or generation according to the structural differences between reference and grayscale frames. Based on this multifunctionality, we propose a novel two-stage coloring strategy. In the first stage, under reference-mask condition, the model autonomously and jointly colors input keyframes in a one-to-many color domain mapping, while temporal coherence constraints are emphasized by modifying the attention mechanism. In the second stage, under reference-guided condition, the model effectively captures the colors of matching structures in the reference, and we further introduce Sliding Reference Grid strategy (SRG) to merge and extract the color features from multiple frames, providing more stable coloring for the grayscale frames. Through this pipeline, we can achieve high-quality and stable video coloring while maintaining the accuracy of detailed colors. Additionally, the two-stage strategy is flexible and detachable, allowing users to adjust the number of selected reference frames to balance coloring quality and efficiency. Extensive experiments demonstrate that our method significantly outperforms previous state-of-the-art models in both qualitative comparison and quantitative measurement.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Color3D: Controllable and Consistent 3D Colorization with Personalized ColorizerYecong Wan, Mingwen Shao, Renlong Wu, Wangmeng ZuoICLR 2026 · 被引用 4 次
- FaceShot: Bring Any Character into LifeJunyao Gao, Yanan Sun, Fei Shen, Xin Jiang 等ICLR 2025
相关 Paper
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li 等ACM MM 2025 · 被引用 2 次
- Versatile Vision Foundation Model for Image and Video ColorizationVukasin Bozic, Abdelaziz Djelouah, Yang Zhang, Radu Timofte 等SIGGRAPH 2024 · 被引用 9 次
- Gray2ColorNet: Transfer More Colors from Reference ImagePeng Lu, Jinbei Yu, Xujun Peng, Zhaoran Zhao 等ACM MM 2020 · 被引用 55 次
- Automatic Controllable Colorization via ImaginationXiaoyan Cong, Yue Wu, Qifeng Chen, Chenyang LeiCVPR 2024
- COCO-LC: Colorfulness Controllable Language-based ColorizationYifan Li, Yuhang Bai, Shuai Yang, Jiaying LiuACM MM 2024 · 被引用 7 次
