Towards High-resolution and Disentangled Reference-based Sketch Colorization
Dingkun Yan, Xinrui Wang, Ru Wang, Zhuoru Li, Jinze Yu, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo
摘要
Sketch colorization models have been widely studied to automate and assist in the creation of animation frames and digital illustrations. However, current methods are still not satisfactory for industrial standard applications in high-resolution synthesis and precise controllability of details. To further enhance the synthesis quality and controllability, we propose an image-referenced sketch colorization method based on the powerful SDXL backbone and leverage sketches as spatial guidance and RGB images as color references. A split cross-attention mechanism is coupled with spatial masks to separately colorize the foreground and background regions to avoid spatial entanglement. A tagger network trained on a massive anime-style image dataset is employed to extract attribution-level information from reference images and integrated into the pipeline to provide precise control signals for synthesis. However, the increased resolution and number of attention layers in the SDXL backbone and precise reference information from the tagger network cause severe entanglement during colorization. We consequently combine a foreground encoder and a background encoder for disentanglement and better synthesis quality. Furthermore, a high-quality annotated and paired sketch colorization dataset is collected for fine-tuning. The proposed method is the first to achieve high resolution high quality sketch colorization with precise control, and obvious outperforms existing methods in quantitative and qualitative validations, as well as user studies in both quality and controllability. Ablation study reveals the influence of each component. Code and dataset will be made publicly available upon paper acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Image Referenced Sketch Colorization Based on Animation Creation WorkflowDingkun Yan, Xinrui Wang, Zhuoru Li, Suguru Saito 等CVPR 2025
- AnimeColor: Reference-based Animation Colorization with Diffusion TransformersYuhong Zhang, Liyao Wang, Han Wang, Danni Wu 等ACM MM 2025 · 被引用 2 次
- Reference-Based Sketch Image Colorization Using Augmented-Self Reference and Dense Semantic CorrespondenceJunsoo Lee, Eungyeup Kim, Yunsung Lee, Dongjun Kim 等CVPR 2020
- Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch InbetweeningJiaming Shen, Kun Hu, Wei Bao, Chang Wen Chen 等ACM MM 2024 · 被引用 5 次
- SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace ColorizationJunpyo Seo, Hanbin Koo, Jieun Yook, Byung-Ro MoonNeurIPS 2025
