CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable Quantization
Yongjie Hu, Yifan Jiang, Ziyun Li, Fei Gao, Henrik Boström, Nannan Wang
Abstract
Face cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get caebd0a5-da32-4505-9eba-069de56de138Related papers
- Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyZhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin et al.AAAI 2022 · 7 citations
- ToonTalker: Cross-Domain Face ReenactmentYuan Gong, Yong Zhang, Xiaodong Cun, Fei Yin et al.ICCV 2023 · 14 citations
- Cross-modal Latent Space Alignment for Image to Avatar TranslationManuel Ladron de Guevara, Yannick Hold-Geoffroy, Jose Echevarria, Cameron Smith et al.ICCV 2023 · 5 citations
- Cartoon Face Recognition: A Benchmark DatasetYi Zheng, Yifan Zhao, Mengyuan Ren, He Yan et al.ACM MM 2020 · 53 citations
- AnyTalk: Multi-modal Driven Multi-domain Talking Head GenerationYu Wang, Yunfei Liu, Fa-Ting Hong, Meng Cao et al.AAAI 2025 · 2 citations
