CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable Quantization
Yongjie Hu, Yifan Jiang, Ziyun Li, Fei Gao, Henrik Boström, Nannan Wang
摘要
Face cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyZhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin 等AAAI 2022 · 被引用 7 次
- ToonTalker: Cross-Domain Face ReenactmentYuan Gong, Yong Zhang, Xiaodong Cun, Fei Yin 等ICCV 2023 · 被引用 14 次
- Cross-modal Latent Space Alignment for Image to Avatar TranslationManuel Ladron de Guevara, Yannick Hold-Geoffroy, Jose Echevarria, Cameron Smith 等ICCV 2023 · 被引用 5 次
- Cartoon Face Recognition: A Benchmark DatasetYi Zheng, Yifan Zhao, Mengyuan Ren, He Yan 等ACM MM 2020 · 被引用 53 次
- AnyTalk: Multi-modal Driven Multi-domain Talking Head GenerationYu Wang, Yunfei Liu, Fa-Ting Hong, Meng Cao 等AAAI 2025 · 被引用 2 次
