CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesis
Dongyu Wang, Dar-Yen Chen, Yi-Zhe Song
Abstract
Sketch-based caricature synthesis suffers from a fundamental failure mode: when identity and shape conditions are combined in diffusion models, they create destructive interference that causes inevitable collapse toward either bland portraits or unrecognizable distortions. We identify the root cause as condition signal contamination -- competing probability distributions in the denoising trajectory that make balanced generation impossible. We present CaricHarmony, the first training-free method that explicitly resolves this contamination through parallel uncontaminated diffusion paths. During inference, we maintain three paths: (pure identity), (pure shape), and (harmonized output). Novel energy functions operating on cross-attention features provide gradient guidance that steers toward optimal balance: ensures sketch fidelity through layout and semantic alignment, while employs token-level correspondence matching robust to extreme distortions. Unlike DemoCaricature requiring 70 seconds per-identity fine-tuning or CaricatureBooth constrained to Bezier curves, CaricHarmony accepts any sketch format and generates in under 16 seconds. Experiments demonstrate state-of-the-art performance: 0.8615 shape CLIP score (vs. 0.8450) under comparable identity consistency score, with 7.81 overall user preference score (vs. 6.06). Our method fundamentally reconceptualizes the ID-shape conflict as conditioning signal contamination for diffusion models, enabling unprecedented creative control while preserving recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da8bde2-5fe8-4372-8a68-def2bf1af973Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- SketchDeco: Training-Free Latent Composition for Precise Sketch ColourisationChaitat Utintu, Yi-Zhe SongCVPR 2026
- CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo BoothZhiyu Qu, Yunqi Miao, Zhensong Zhang, Jifei Song et al.CVPR 2025
- One-Shot Reference-based Structure-Aware Image to Sketch SynthesisRui Yang, Honghong Yang, Li Zhao, Qin Lei et al.AAAI 2025 · 2 citations
- FreePIH: Training-Free Painterly Image Harmonization with Diffusion ModelRuibin Li, Jingcai Guo, Qihua Zhou, Song GuoACM MM 2024 · 2 citations
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu et al.CVPR 2026 · 3 citations
