Hashing Neural Video Decomposition with Multiplicative Residuals in Space-Time
Cheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun, Hwann-Tzong Chen
Abstract
We present a video decomposition method that facilitates layer-based editing of videos with spatiotemporally varying lighting and motion effects. Our neural model decomposes an input video into multiple layered representations, each comprising a 2D texture map, a mask for the original video, and a multiplicative residual characterizing the spatiotemporal variations in lighting conditions. A single edit on the texture maps can be propagated to the corresponding locations in the entire video frames while preserving other contents’ consistencies. Our method efficiently learns the layer-based neural representations of a 1080p video in 25s per frame via coordinate hashing and allows real-time rendering of the edited result at 71 fps on a single GPU. Qualitatively, we run our method on various videos to show its effectiveness in generating high-quality editing effects. Quantitatively, we propose to adopt feature-tracking evaluation metrics for objectively assessing the consistency of video editing. Project page: https://lightbulb12294.github.io/hashing-nvd/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Splatter a Video: Video Gaussian Representation for Versatile ProcessingYang-Tian Sun, Yihua Huang, Lin Ma, Xiaoyang Lyu et al.NeurIPS 2024 · 41 citations
- NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video EditingTing-Hsuan Chen, Jiewen Chan, Hau-Shiang Shiu, Shih-Han Yen et al.NeurIPS 2024 · 8 citations
- HyperNVD: Accelerating Neural Video Decomposition via HypernetworksMaria Pilligua, Danna Xue, Javier Vazquez-CorralCVPR 2025
Builds on10
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- GAN-Supervised Dense Visual AlignmentWilliam S. Peebles, Jun-Yan Zhu, Richard Zhang, Antonio Torralba et al.CVPR 2022 · 50 citations
- Deformable Sprites for Unsupervised Video DecompositionVickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa et al.CVPR 2022 · 45 citations
- Learning Pixel Trajectories with Multiscale Contrastive Random WalksZhangxing Bian, Allan Jabri, Alexei A. Efros, Andrew OwensCVPR 2022 · 35 citations
Related papers
- Video Decomposition Prior: Editing Videos Layer by LayerGaurav Shrivastava, Ser-Nam Lim, Abhinav ShrivastavaICLR 2024 · 11 citations
- Editable free-viewpoint video using a layered neural representationJiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao et al.SIGGRAPH 2021 · 80 citations
- RT-VENet: A Convolutional Network for Real-time Video EnhancementMohan Zhang, Qiqi Gao, Jinglu Wang, Henrik Turbell et al.ACM MM 2020 · 5 citations
- STRIVE: Scene Text Replacement In VideosVijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart et al.ICCV 2021 · 14 citations
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 219 citations
