Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency
Zhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin, Dandan Tu, Shuhui Wang, Zheng-Jun Zha
Abstract
In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the style effect of generated images, video cartoonization has additional requirements on the temporal consistency. In this paper, we propose a spatially-adaptive semantic alignment framework with perceptual motion consistency for coherent video cartoonization in an unsupervised manner. The semantic alignment module is designed to restore deformation of semantic structure caused by spatial information lost in the encoder-decoder architecture. Furthermore, we devise the spatio-temporal correlative map as a style-independent, global-aware regularization on the perceptual motion consistency. Deriving from similarity measurement of high-level features in photo and cartoon frames, it captures global semantic information beyond raw pixel-value in optical flow. Besides, the similarity measurement disentangles temporal relationships from domain-specific style properties, which helps regularize the temporal consistency without hurting style effects of cartoon images. Qualitative and quantitative experiments demonstrate our method is able to generate highly stylistic and temporal consistent cartoon videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d4c8e1-e0de-4566-a6bb-9a8723b5ce44Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang et al.AAAI 2021 · 197 citations
- Blind Video Temporal Consistency via Deep Video PriorChenyang Lei, Yazhou Xing, Qifeng ChenNeurIPS 2020 · 134 citations
- Consistent Video Style Transfer via Compound RegularizationWenjing Wang, Jizheng Xu, Li Zhang, Yue Wang et al.AAAI 2020 · 50 citations
- Multimodal Structure-Consistent Image-to-Image TranslationChe-Tsung Lin, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong LaiAAAI 2020 · 24 citations
- Learning to Transfer: Unsupervised Domain Translation via Meta-LearningJianxin Lin, Yijun Wang, Zhibo Chen, Tianyu HeAAAI 2020 · 10 citations
Related papers
- Preserving Global and Local Temporal Consistency for Arbitrary Video Style TransferXinxiao Wu, Jialu ChenACM MM 2020 · 14 citations
- CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable QuantizationYongjie Hu, Yifan Jiang, Ziyun Li, Fei Gao et al.ACM MM 2025
- Learning Temporally and Semantically Consistent Unpaired Video-to-Video Translation through Pseudo-Supervision from Synthetic Optical FlowKaihong Wang, Kumar Akash, Teruhisa MisuAAAI 2022 · 16 citations
- Cartoon-Flow: A Flow-Based Generative Adversarial Network for Arbitrary-Style Photo CartoonizationJieun Lee, Hyeonwoo Kim, Jonghwa Shim, Eenjun HwangACM MM 2022 · 14 citations
- STRIVE: Scene Text Replacement In VideosVijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart et al.ICCV 2021 · 14 citations
