S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window Attention
Chiyu Zhang, Xiaogang Xu, Lei Wang, Zaiyan Dai, Jun Yang
Abstract
Transformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a novel hierarchical vision transformer designed for style transfer. S2WAT employs attention computation in diverse window shapes to capture both short- and long-range dependencies. The merged dependencies utilize the "Attn Merge" strategy, which adaptively determines spatial weights based on their relevance to the target. Extensive experiments on representative datasets show the proposed method's effectiveness compared to state-of-the-art (SOTA) transformer-based and other approaches. The code and pre-trained models are available at https://github.com/AlienZhang1996/S2WAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a837752-8e7b-4113-9f34-a47699797554Cited by top-tier papers6
- Styl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and StylesPeng Wang, Xiang Liu, Peidong LiuNeurIPS 2025 · 8 citations
- SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHongda Liu, Longguang Wang, Ye Zhang, Ziru Yu et al.CVPR 2025
- SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferChunnan Shang, Zhizhong Wang, Hongwei Wang, Xiangming MengCVPR 2025
- StyleFM: Frequency Manipulation Empowered by Recursive Attention on Diffusion Models for Arbitrary Style TransferYingnan Ma, Zhenye Liu, Siying Liu, Anup BasuAAAI 2026
- Z*: Zero-shot Style Transfer via Attention ReweightingYingying Deng, Xiangyu He, Fan Tang, Weiming DongCVPR 2024
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
Related papers
- StyTr2: Image Style Transfer with TransformersYingying Deng, Fan Tang, Weiming Dong, Chongyang Ma et al.CVPR 2022 · 345 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- Activating More Pixels in Image Super-Resolution TransformerXiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao et al.CVPR 2023
- Styleformer: Transformer based Generative Adversarial Networks with Style VectorJeeseung Park, Younggeun KimCVPR 2022 · 49 citations
- Learning Spatial Decay for Vision TransformersYuxin Mao, Zhen Qin, Jinxing Zhou, Bin Fan et al.AAAI 2026 · 1 citation
