StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
Boyu He, Yunfan Ye, Chang Liu, Weishang Wu, Fang Liu, Zhiping Cai
Abstract
Despite the advancements in diffusion-based image style transfer, existing methods are commonly limited by 1) semantic gap: the style reference could miss proper content semantics, causing uncontrollable stylization; 2) reliance on extra constraints (e.g., semantic masks) restricting applicability; 3) rigid feature associations lacking adaptive global-local alignment, failing to balance fine-grained stylization and global content preservation. These limitations, particularly the inability to flexibly leverage style inputs, fundamentally restrict style transfer in terms of personalization, accuracy, and adaptability. To address these, we propose StyleGallery, a training-free and semantic-aware framework that supports arbitrary reference images as input and enables effective personalized customization. It comprises three core stages: semantic region segmentation (adaptive clustering on latent diffusion features to divide regions without extra inputs); clustered region matching (block filtering on extracted features for precise alignment); and style transfer optimization (energy function-guided diffusion sampling with regional style loss to optimize stylization). Experiments on our introduced benchmark demonstrate that StyleGallery outperforms state-of-the-art methods in content structure preservation, regional stylization, interpretability, and personalized customization, particularly when leveraging multiple style references.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c99b1d68-0e52-4e31-94a4-548f6268ddb0Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li et al.ICCV 2021 · 421 citations
Related papers
- StyleFM: Frequency Manipulation Empowered by Recursive Attention on Diffusion Models for Arbitrary Style TransferYingnan Ma, Zhenye Liu, Siying Liu, Anup BasuAAAI 2026
- Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion ModelsRui Jiang, Xinghe Fu, Guangcong Zheng, Teng Li et al.AAAI 2025 · 2 citations
- CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style TransferWenbo Nie, Zixiang Li, Renshuai Tao, Bin WU et al.ICLR 2026 · 2 citations
- RB-Modulation: Training-Free Stylization using Reference-Based ModulationLitu Rout, Yujia Chen, Nataniel Ruiz, Abhishek Kumar et al.ICLR 2025
- StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferRuojun Xu, Weijie Xi, Xiaodi Wang, Yongbo Mao et al.CVPR 2025
