PNeRV: Enhancing Spatial Consistency via Pyramidal Neural Representation for Videos
Qi Zhao, M. Salman Asif, Zhan Ma
Abstract
The primary focus of Neural Representation for Videos (NeRV) is to effectively model its spatiotemporal consis-tency. However, current NeRV systems often face a signif-icant issue of spatial inconsistency, leading to decreased perceptual quality. To address this issue, we introduce the Pyramidal Neural Representation for Videos (PNeRV), which is built on a multi-scale information connection and comprises a lightweight rescaling operator, Kronecker Fully-connected layer (KFc), and a Benign Selective Mem-ory (BSM) mechanism. The KFc, inspired by the tensor de-composition of the vanilla Fully-connected layer, facilitates low-cost rescaling and global correlation modeling. BSM merges high-level features with granular ones adaptively. Furthermore, we provide an analysis based on the Univer-sal Approximation Theory of the NeRV system and vali-date the effectiveness of the proposed PNeRV. We conducted comprehensive experiments to demonstrate that PNeRV sur-passes the performance of contemporary NeRV models, achieving the best results in video regression on UVG and DAVIS under various metrics (PSNR, SSIM, LPIPS, and FVD). Compared to vanilla N eRV, P N eRV achieves <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> dB gain in PSNR and a 231% increase in FVD on UVG, along with <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> dB PSNR and 634% FVD increase on DAVIS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Tree-NeRV: Efficient Non-Uniform Sampling for Neural Video Representation via Tree-Structured Feature GridsJiancheng Zhao, Yifan Zhan, Qingtian Zhu, Mingze Ma et al.ICCV 2025 · 4 citations
- Exploring State-Space Models for Data-Specific Neural RepresentationsJinsung Lee, Suha KwakICLR 2026
- On Quantizing Neural Representation for Variable-Rate Video CodingJunqi Shi, Zhujia Chen, Hanfei Li, Qi Zhao et al.ICLR 2025
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- HNeRV: A Hybrid Neural Representation for VideosHao Chen, Matthew Gwilliam, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2023
- DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic CodesHao Yan, Zhihui Ke, Xiaobo Zhou, Tie Qiu et al.CVPR 2024 · 18 citations
- DNeRV: Modeling Inherent Dynamics via Difference Neural Representation for VideosQi Zhao, M. Salman Asif, Zhan MaCVPR 2023
- MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal GuidanceJialong Guo, Ke Liu, Jiangchao Yao, Zhihua Wang et al.AAAI 2025 · 7 citations
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
