NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-Wise Modeling
Shishira R. Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang, Kwot Sin Lee, Patrick Poirson, Pengxiang Wu, Chen Wang, Abhinav Shrivastava
Abstract
Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architectures which do not scale to longer videos or higher resolutions. To address these issues, we propose NIRVANA, which treats videos as groups of frames and fits separate networks to each group performing patch-wise prediction. The video representation is modeled autoregressively, with networks fit on a current group initialized using weights from the previous group's model. To further enhance efficiency, we perform quantization of the network parameters during training, requiring no post-hoc pruning or quantization. When compared with previous works on the benchmark UVG dataset, NIRVANA improves encoding quality from 37.36 to 37.70 (in terms of PSNR) and the encoding speed by 12×, while maintaining the same compression rate. In contrast to prior video INR works which struggle with larger resolution and longer videos, we show that our algorithm is highly flexible and scales naturally due to its patch-wise and autoregressive designs. Moreover, our method achieves variable bitrate compression by adapting to videos with varying inter-frame motion. NIR-VANA achieves 6× decoding speed and scales well with more GPUs, making it practical for various deployment scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 086a92f1-ca19-4d73-9334-5cd7e7cf0092Cited by top-tier papers5
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2024 · 44 citations
- NTK-Guided Implicit Neural TeachingChen Zhang, Wei Zuo, Bingyang Cheng, Yikun Wang et al.CVPR 2026 · 3 citations
- Good, Cheap, and Fast: Overfitted Image Compression with Wasserstein DistortionJona Ballé, Luca Versari, Emilien Dupont, Hyunjik Kim et al.CVPR 2025
- Combining Frame and GOP Embeddings for Neural Video RepresentationJens Eirik Saethre, Roberto Azevedo, Christopher SchroersCVPR 2024
- EVOS: Efficient Implicit Neural Training via EVOlutionary SelectorWeixiang Zhang, Shuzhao Xie, Chengwei Ren, Siyi Xie et al.CVPR 2025
Builds on17
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
- Training with Quantization Noise for Extreme Model CompressionPierre Stock, Angela Fan, Benjamin Graham, Edouard Grave et al.ICLR 2021 · 262 citations
Related papers
- Video Compression with Entropy-Constrained Neural RepresentationsCarlos Gomes, Roberto Azevedo, Christopher SchroersCVPR 2023
- Towards Scalable Neural Representation for Diverse VideosBo He, Xitong Yang, Hanyu Wang, Zuxuan Wu et al.CVPR 2023
- HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2023 · 132 citations
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 20 citations
- DNeRV: Modeling Inherent Dynamics via Difference Neural Representation for VideosQi Zhao, M. Salman Asif, Zhan MaCVPR 2023
