Two Birds, One Stone: A Unified Framework for Joint Learning of Image and Video Style Transfers
Bohai Gu, Heng Fan, Libo Zhang
摘要
Current arbitrary style transfer models are limited to either image or video domains. In order to achieve satisfying image and video style transfers, two different models are inevitably required with separate training processes on image and video domains, respectively. In this paper, we show that this can be precluded by introducing UniST, a Unified Style Transfer framework for both images and videos. At the core of UniST is a domain interaction transformer (DIT ), which first explores context information within the specific domain and then interacts contextualized domain information for joint learning. In particular, DIT enables exploration of temporal information from videos for the image style transfer task and meanwhile allows rich appearance texture from images for video style transfer, thus leading to mutual benefits. Considering heavy computation of traditional multi-head self-attention, we present a simple yet effective axial multi-head self-attention (AMSA) for DIT , which improves computational efficiency while maintains style transfer performance. To verify the effectiveness of UniST, we conduct extensive experiments on both image and video style transfer tasks and show that UniST performs favorably against state-of-the-art approaches on both tasks. Code is available at https://github.com/NevSNev/UniST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerXing Cui, Peipei Li, Zekun Li, Xuannan Liu 等NeurIPS 2024 · 被引用 11 次
- CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingKornel Howil, Joanna Waczynska, Piotr Borycki, Tadeusz Dziarmaga 等NeurIPS 2025 · 被引用 11 次
- Video Color Grading via Look-Up Table GenerationSeunghyun Shin, Dongmin Shin, Jisu Shin, Hae-Gon Jeon 等ICCV 2025 · 被引用 2 次
- Semantix: An Energy-guided Sampler for Semantic Style TransferHuiang He, Minghui Hu, Chuanxia Zheng, Chaoyue Wang 等ICLR 2025
- FlowStyler: Artistic Video Stylization Via Transformation Fields TransportsYuning Gong, Jiaming Chen, Xiaohua Ren, Yuanjun Liao 等ICCV 2025
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Incorporating Convolution Designs into Visual TransformersKun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou 等ICCV 2021 · 被引用 581 次
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li 等ICCV 2021 · 被引用 421 次
- Artistic Style Transfer with Internal-external Learning and Contrastive LearningHaibo Chen, Lei Zhao, Zhizhong Wang, Huiming Zhang 等NeurIPS 2021 · 被引用 243 次
相关 Paper
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos 等ACM MM 2025 · 被引用 2 次
- Dual-head Genre-instance Transformer Network for Arbitrary Style TransferMeichen Liu, Shuting He, Songnan Lin, Bihan WenACM MM 2024 · 被引用 3 次
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang 等ICLR 2026 · 被引用 90 次
- UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesChen Tang, Xinzhu Ma, Encheng Su, Xiufeng Song 等CVPR 2025
- UNIST: Unpaired Neural Implicit Shape Translation NetworkQimin Chen, Johannes Merz, Aditya Sanghi, Hooman Shayani 等CVPR 2022 · 被引用 15 次
