Cobra: Efficient Line Art COlorization with BRoAder References
Junhao Zhuang, Lingen Li, Xuan Ju, Zhaoyang Zhang, Chun Yuan, Ying Shan
摘要
The comic production industry requires reference-based line art colorization with high accuracy, efficiency, contextual consistency, and flexible control.A comic page often involves diverse characters, objects, and backgrounds, which complicates the coloring process. Despite advancements in diffusion models for image generation, their application in line art colorization remains limited, facing challenges related to handling extensive reference images, time-consuming inference, and flexible control. We investigate the necessity of extensive contextual image guidance on the quality of line art colorization. To address these challenges, we introduce Cobra, an efficient and versatile method that supports color hints and utilizes over 200 reference images while maintaining low latency. Central to Cobra is a Causal Sparse DiT architecture, which leverages specially designed positional encodings, causal sparse attention, and Key-Value Cache to effectively manage long-context references and ensure color identity consistency. Results demonstrate that Cobra achieves accurate line art colorization through extensive contextual reference, significantly enhancing inference speed and interactivity, thereby meeting critical industrial demands. We release our codes and models on our project page: https://zhuang2002.github.io/Cobra/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ToonComposer: Streamlining Cartoon Production with Generative Post-KeyframingLingen Li, Guangzhi Wang, Zhaoyang Zhang, Yaowei Li 等ICLR 2026 · 被引用 8 次
- Towards High-resolution and Disentangled Reference-based Sketch ColorizationDingkun Yan, Xinrui Wang, Ru Wang, Zhuoru Li 等CVPR 2026 · 被引用 1 次
- SketchDeco: Training-Free Latent Composition for Precise Sketch ColourisationChaitat Utintu, Yi-Zhe SongCVPR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
相关 Paper
- AnimeColor: Reference-based Animation Colorization with Diffusion TransformersYuhong Zhang, Liyao Wang, Han Wang, Danni Wu 等ACM MM 2025 · 被引用 2 次
- MangaNinja: Line Art Colorization with Precise Reference FollowingZhiheng Liu, Ka Leong Cheng, Xi Chen, Jie Xiao 等CVPR 2025
- AniDoc: Animation Creation Made EasierYihao Meng, Hao Ouyang, Hanlin Wang, Qiuyu Wang 等CVPR 2025
- Magiccolor: Multi-Instance Sketch ColorizationYinhan Zhang, Yue Ma, Bingyuan Wang, Qifeng Chen 等ICCV 2025 · 被引用 3 次
- Image Referenced Sketch Colorization Based on Animation Creation WorkflowDingkun Yan, Xinrui Wang, Zhuoru Li, Suguru Saito 等CVPR 2025
