Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion Network
Sizhe Zheng, Pan Gao, Peng Zhou, Jie Qin
Abstract
Style transfer aims to render an image with the artistic features of a style image, while maintaining the origi-nal structure. Various methods have been put forward for this task, but some challenges still exist. For instance, it is difficult for CNN-based methods to handle global information and long-range dependencies between input images, for which transformer-based methods have been proposed. Although transformers can better model the relationship between content and style images, they require high-cost hard-ware and time-consuming inference. To address these is-sues, we design a novel transformer model that includes only the encoder, thus significantly reducing the computational cost. In addition, we also find that existing style transfer methods may lead to images under-stylied or missing content. In order to achieve better stylization, we de-sign a content feature extractor and a style feature extrac-tor, based on which pure content and style images can be fed to the transformer. Finally, we propose a novel network termed Puff-Net, i.e., pure content and style feature fusion network. Through qualitative and quantitative experiments, we demonstrate the advantages of our model compared to state-of-the-art ones in the literature. The code is available at https://github.com/ZszYmy9/Puff-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- AnimateQR: Bridging Aesthetics and Functionality in Dynamic QR Code GenerationGuangyang Wu, Huayu Zheng, Siqi Luo, Guangtao Zhai et al.NeurIPS 2025 · 2 citations
- SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHongda Liu, Longguang Wang, Ye Zhang, Ziru Yu et al.CVPR 2025
- FEAT: Fashion Editing and Try-On from Any DesignSoye Kwon, Keonyoung Lee, Dahuin Jung, Jaekoo LeeCVPR 2026
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- StyTr2: Image Style Transfer with TransformersYingying Deng, Fan Tang, Weiming Dong, Chongyang Ma et al.CVPR 2022 · 345 citations
- StyleFormer: Real-time Arbitrary Style Transfer via Parametric Style CompositionXiaolei Wu, Zhihao Hu, Lu Sheng, Dong XuICCV 2021 · 130 citations
- Z*: Zero-shot Style Transfer via Attention ReweightingYingying Deng, Xiangyu He, Fan Tang, Weiming DongCVPR 2024
- Dual-head Genre-instance Transformer Network for Arbitrary Style TransferMeichen Liu, Shuting He, Songnan Lin, Bihan WenACM MM 2024 · 3 citations
- Handwriting TransformersAnkan Kumar Bhunia, Salman H. Khan, Hisham Cholakkal, Rao Muhammad Anwer et al.ICCV 2021 · 64 citations
