Two Birds, One Stone: A Unified Framework for Joint Learning of Image and Video Style Transfers
Bohai Gu, Heng Fan, Libo Zhang
Abstract
Current arbitrary style transfer models are limited to either image or video domains. In order to achieve satisfying image and video style transfers, two different models are inevitably required with separate training processes on image and video domains, respectively. In this paper, we show that this can be precluded by introducing UniST, a Unified Style Transfer framework for both images and videos. At the core of UniST is a domain interaction transformer (DIT ), which first explores context information within the specific domain and then interacts contextualized domain information for joint learning. In particular, DIT enables exploration of temporal information from videos for the image style transfer task and meanwhile allows rich appearance texture from images for video style transfer, thus leading to mutual benefits. Considering heavy computation of traditional multi-head self-attention, we present a simple yet effective axial multi-head self-attention (AMSA) for DIT , which improves computational efficiency while maintains style transfer performance. To verify the effectiveness of UniST, we conduct extensive experiments on both image and video style transfer tasks and show that UniST performs favorably against state-of-the-art approaches on both tasks. Code is available at https://github.com/NevSNev/UniST.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ec9ceb9-3ff9-4f1d-85e9-55de3ac354e2Cited by top-tier papers5
- Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerXing Cui, Peipei Li, Zekun Li, Xuannan Liu et al.NeurIPS 2024 · 11 citations
- CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingKornel Howil, Joanna Waczynska, Piotr Borycki, Tadeusz Dziarmaga et al.NeurIPS 2025 · 11 citations
- Video Color Grading via Look-Up Table GenerationSeunghyun Shin, Dongmin Shin, Jisu Shin, Hae-Gon Jeon et al.ICCV 2025 · 2 citations
- Semantix: An Energy-guided Sampler for Semantic Style TransferHuiang He, Minghui Hu, Chuanxia Zheng, Chaoyue Wang et al.ICLR 2025
- FlowStyler: Artistic Video Stylization Via Transformation Fields TransportsYuning Gong, Jiaming Chen, Xiaohua Ren, Yuanjun Liao et al.ICCV 2025
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Incorporating Convolution Designs into Visual TransformersKun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou et al.ICCV 2021 · 581 citations
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li et al.ICCV 2021 · 421 citations
- Artistic Style Transfer with Internal-external Learning and Contrastive LearningHaibo Chen, Lei Zhao, Zhizhong Wang, Huiming Zhang et al.NeurIPS 2021 · 243 citations
Related papers
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos et al.ACM MM 2025 · 2 citations
- Dual-head Genre-instance Transformer Network for Arbitrary Style TransferMeichen Liu, Shuting He, Songnan Lin, Bihan WenACM MM 2024 · 3 citations
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang et al.ICLR 2026 · 90 citations
- UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesChen Tang, Xinzhu Ma, Encheng Su, Xiufeng Song et al.CVPR 2025
- UNIST: Unpaired Neural Implicit Shape Translation NetworkQimin Chen, Johannes Merz, Aditya Sanghi, Hooman Shayani et al.CVPR 2022 · 15 citations
