HyperNVD: Accelerating Neural Video Decomposition via Hypernetworks
Maria Pilligua, Danna Xue, Javier Vazquez-Corral
Abstract
Decomposing a video into a layer-based representation is crucial for easy video editing for the creative industries, as it enables independent editing of specific layers. Existing video-layer decomposition models rely on implicit neural representations (INRs) trained independently for each video, making the process time-consuming when applied to new videos. Noticing this limitation, we propose a metalearning strategy to learn a generic video decomposition model to speed up the training on new videos. Our model is based on a hypernetwork architecture which, given a videoencoder embedding, generates the parameters for a compact INR-based neural video decomposition model. Our strategy mitigates the problem of single-video overfitting and, importantly, shortens the convergence of video decomposition on new, unseen videos. Our code is available at: https://hypernvd.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
Related papers
- MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal GuidanceJialong Guo, Ke Liu, Jiangchao Yao, Zhihua Wang et al.AAAI 2025 · 7 citations
- Generalizable Implicit Neural Representations via Instance Pattern ComposersChiheon Kim, Doyup Lee, Saehoon Kim, Minsu Cho et al.CVPR 2023
- Video Compression with Entropy-Constrained Neural RepresentationsCarlos Gomes, Roberto Azevedo, Christopher SchroersCVPR 2023
- NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-Wise ModelingShishira R. Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang et al.CVPR 2023
- Hashing Neural Video Decomposition with Multiplicative Residuals in Space-TimeCheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun, Hwann-Tzong ChenICCV 2023 · 5 citations
