ZePo: Zero-Shot Portrait Stylization with Faster Sampling
Jin Liu, Huaibo Huang, Jie Cao, Ran He
Abstract
Diffusion-based text-to-image generation models have significantly advanced the field of art content synthesis. However, current portrait stylization methods generally require either model fine-tuning based on examples or the employment of DDIM Inversion to revert images to noise space, both of which substantially decelerate the image generation process. To overcome these limitations, this paper presents an inversion-free portrait stylization framework based on diffusion models that accomplishes content and style feature fusion in merely four sampling steps. We observed that Latent Consistency Models employing consistency distillation can effectively extract representative Consistency Features from noisy images. To blend the Consistency Features extracted from both content and style images, we introduce a Style Enhancement Attention Control technique that meticulously merges content and style features within the attention space of the target image. Moreover, we propose a feature merging strategy to amalgamate redundant features in Consistency Features, thereby reducing the computational load of attention control. Extensive experiments have validated the effectiveness of our proposed framework in enhancing stylization efficiency and fidelity. The code is available at https://github.com/liujin112/ZePo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd1baeed-ed20-4ff5-b837-b703cdc878d4Cited by top-tier papers1
Ask how each one uses itBuilds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Tuning-Free Inversion-Enhanced Control for Consistent Image EditingXiaoyue Duan, Shuhao Cui, Guoliang Kang, Baochang Zhang et al.AAAI 2024 · 13 citations
- Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 StepsNikita Starodubcev, Mikhail Khoroshikh, Artem Babenko, Dmitry BaranchukNeurIPS 2024 · 18 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- Morse: Dual-Sampling for Lossless Acceleration of Diffusion ModelsChao Li, Jiawei Fan, Anbang YaoICML 2025
- Inversion-Free Image Editing with Language-Guided Diffusion ModelsSihan Xu, Yidong Huang, Jiayi Pan, Ziqiao Ma et al.CVPR 2024 · 12 citations
