ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
Lingen Li, Guangzhi Wang, Zhaoyang Zhang, Yaowei Li, Xiaoyu Li, Qi Dou, Jinwei Gu, Tianfan Xue, Ying Shan
Abstract
Traditional cartoon and anime production involves keyframing, inbetweening, and colorization stages, which require intensive manual effort. Despite recent advances in AI, existing methods often handle these stages separately, leading to error accumulation and artifacts. For instance, inbetweening approaches struggle with large motions, while colorization methods require dense per-frame sketches. To address this, we introduce ToonComposer, a generative model that unifies inbetweening and colorization into a single post-keyframing stage. ToonComposer employs a sparse sketch injection mechanism to provide precise control using keyframe sketches. Additionally, we propose a novel cartoon adaptation method with the spatial low-rank adapter to effectively tailor a modern video foundation model to the cartoon domain while keeping its temporal prior intact. Requiring as few as a single sketch and a colored reference frame, ToonComposer excels with sparse inputs, while also supporting multiple sketches at any temporal location for more precise motion control. This dual capability reduces manual workload and improves flexibility, empowering artists in real-world scenarios. To evaluate our model, we further created PKBench, a benchmark featuring human-drawn sketches that simulate real-world use cases. Our evaluation demonstrates that ToonComposer outperforms existing methods in visual quality, motion consistency, and production efficiency, offering a superior and more flexible solution for AI-assisted cartoon production.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd0c7485-a8de-4434-ac44-cac3aed0973eCited by top-tier papers3
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang et al.CVPR 2026 · 10 citations
- Meta-CoT: Enhancing Granularity and Generalization in Image EditingShiyi Zhang, Yiji Cheng, Tiankai Hang, Zijin Yin et al.CVPR 2026 · 3 citations
- AniMimic: Imitating 3D Animation from Video PriorsTianyi Xie, Yunuo Chen, Yaowei Guo, Yin Yang et al.CVPR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- AniDoc: Animation Creation Made EasierYihao Meng, Hao Ouyang, Hanlin Wang, Qiuyu Wang et al.CVPR 2025
- AnimeColor: Reference-based Animation Colorization with Diffusion TransformersYuhong Zhang, Liyao Wang, Han Wang, Danni Wu et al.ACM MM 2025 · 2 citations
- Image Referenced Sketch Colorization Based on Animation Creation WorkflowDingkun Yan, Xinrui Wang, Zhuoru Li, Suguru Saito et al.CVPR 2025
- Adaptive Interpolation-Synthesis for Motion In-Betweening on Keyframe-Based AnimationAnton Raël, Julien Boucher, Antoine LhermitteSIGGRAPH 2026
- DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference ImagesKazuma Nagata, Naoshi KanekoICCV 2025 · 1 citation
