DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images
Kazuma Nagata, Naoshi Kaneko
Abstract
Automatic colorization of line drawings has been widely studied to reduce the labor cost of hand-drawn anime production. Deep learning approaches, including image/video generation and feature-based correspondence, have improved accuracy but struggle with occlusions, pose variations, and viewpoint changes. To address these challenges, we propose DACoN, a framework that leverages foundation models to capture part-level semantics, even in line drawings. Our method fuses low-resolution semantic features from foundation models with high-resolution spatial features from CNNs for fine-grained yet robust feature extraction. In contrast to previous methods that rely on the Multiplex Transformer and support only one or two reference images, DACoN removes this constraint, allowing any number of references. Quantitative and qualitative evaluations demonstrate the benefits of using multiple reference images, achieving superior colorization performance. Our code and model are available at https://github.com/kzmngt/DACoN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 099d0010-cd5f-44cd-b6f7-839ae5805bf6Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera et al.NeurIPS 2023 · 371 citations
- The Animation Transformer: Visual Correspondence via Segment MatchingEvan Casey, Víctor Pérez, Zhuoru LiICCV 2021 · 41 citations
Related papers
- Region-Wise Correspondence Prediction between Manga Line Art ImagesYingxuan Li, Jiafeng Mao, Qianru Qiu, Yusuke MatsuiCVPR 2026
- AnimeColor: Reference-based Animation Colorization with Diffusion TransformersYuhong Zhang, Liyao Wang, Han Wang, Danni Wu et al.ACM MM 2025 · 2 citations
- AniDoc: Animation Creation Made EasierYihao Meng, Hao Ouyang, Hanlin Wang, Qiuyu Wang et al.CVPR 2025
- Tag2Pix: Line Art Colorization Using Text Tag With SECat and Changing LossHyunsu Kim, Ho Young Jhoo, Eunhyeok Park, Sungjoo YooICCV 2019 · 119 citations
- A Unified Framework for Industrial Cel-Animation Colorization with Temporal-Structural AwarenessXiaoyi Feng, Tao Huang, Peng Wang, Zizhou Huang et al.ICCV 2025 · 3 citations
