ViSt3D: Video Stylization with 3D CNN
Ayush Pande, Gaurav Sharma
摘要
Visual stylization has been a very popular research area in recent times. While image stylization has seen a rapid advancement in the recent past, video stylization, while being more challenging, is relatively less explored. The immediate method of stylizing videos by stylizing each frame independently has been tried with some success. To the best of our knowledge, we present the first approach to video stylization using 3D CNN directly, building upon insights from 2D image stylization. Stylizing video is highly challenging, as the appearance and video motion, which includes both camera and subject motions, are inherently entangled in the representations learnt by a 3D CNN. Hence, a naive extension of 2D CNN stylization methods to 3D CNN does not work. To perform stylization with 3D CNN, we propose to explicitly disentangle motion and appearance, stylize the appearance part, and then add back the motion component and decode the final stylized video. In addition, we propose a dataset, curated from existing datasets, to train video stylization networks. We also provide an independently collected test set to study the generalization of video stylization methods. We provide results on this test dataset comparing the proposed method with 2D stylization methods applied frame by frame. We show successful stylization with 3D CNN for the first time, and obtain better stylization in terms of texture cf. the existing 2D methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingKornel Howil, Joanna Waczynska, Piotr Borycki, Tadeusz Dziarmaga 等NeurIPS 2025 · 被引用 11 次
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 被引用 7 次
它引用的顶会 Paper3
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li 等ICCV 2021 · 被引用 421 次
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang 等AAAI 2021 · 被引用 197 次
- Arbitrary Style Transfer via Multi-Adaptation NetworkYingying Deng, Fan Tang, Weiming Dong, Wen Sun 等ACM MM 2020 · 被引用 194 次
相关 Paper
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 被引用 109 次
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
- PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style MappingJiafu Chen, Wei Xing, Jiakai Sun, Tianyi Chu 等AAAI 2024 · 被引用 2 次
- PV3D: A 3D Generative Model for Portrait Video GenerationEric Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Wenqing Zhang 等ICLR 2023 · 被引用 3 次
- G3AN: Disentangling Appearance and Motion for Video GenerationYaohui Wang, Piotr Bilinski, François Brémond, Antitza DantchevaCVPR 2020
