ViSt3D: Video Stylization with 3D CNN
Ayush Pande, Gaurav Sharma
Abstract
Visual stylization has been a very popular research area in recent times. While image stylization has seen a rapid advancement in the recent past, video stylization, while being more challenging, is relatively less explored. The immediate method of stylizing videos by stylizing each frame independently has been tried with some success. To the best of our knowledge, we present the first approach to video stylization using 3D CNN directly, building upon insights from 2D image stylization. Stylizing video is highly challenging, as the appearance and video motion, which includes both camera and subject motions, are inherently entangled in the representations learnt by a 3D CNN. Hence, a naive extension of 2D CNN stylization methods to 3D CNN does not work. To perform stylization with 3D CNN, we propose to explicitly disentangle motion and appearance, stylize the appearance part, and then add back the motion component and decode the final stylized video. In addition, we propose a dataset, curated from existing datasets, to train video stylization networks. We also provide an independently collected test set to study the generalization of video stylization methods. We provide results on this test dataset comparing the proposed method with 2D stylization methods applied frame by frame. We show successful stylization with 3D CNN for the first time, and obtain better stylization in terms of texture cf. the existing 2D methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5ca5cae-78d8-4dc4-beba-fb6997d22399Cited by top-tier papers2
- CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingKornel Howil, Joanna Waczynska, Piotr Borycki, Tadeusz Dziarmaga et al.NeurIPS 2025 · 11 citations
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 7 citations
Builds on3
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li et al.ICCV 2021 · 421 citations
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang et al.AAAI 2021 · 197 citations
- Arbitrary Style Transfer via Multi-Adaptation NetworkYingying Deng, Fan Tang, Weiming Dong, Wen Sun et al.ACM MM 2020 · 194 citations
Related papers
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 109 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style MappingJiafu Chen, Wei Xing, Jiakai Sun, Tianyi Chu et al.AAAI 2024 · 2 citations
- PV3D: A 3D Generative Model for Portrait Video GenerationEric Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Wenqing Zhang et al.ICLR 2023 · 3 citations
- G3AN: Disentangling Appearance and Motion for Video GenerationYaohui Wang, Piotr Bilinski, François Brémond, Antitza DantchevaCVPR 2020
