PNeRV: Enhancing Spatial Consistency via Pyramidal Neural Representation for Videos
Qi Zhao, M. Salman Asif, Zhan Ma
摘要
The primary focus of Neural Representation for Videos (NeRV) is to effectively model its spatiotemporal consis-tency. However, current NeRV systems often face a signif-icant issue of spatial inconsistency, leading to decreased perceptual quality. To address this issue, we introduce the Pyramidal Neural Representation for Videos (PNeRV), which is built on a multi-scale information connection and comprises a lightweight rescaling operator, Kronecker Fully-connected layer (KFc), and a Benign Selective Mem-ory (BSM) mechanism. The KFc, inspired by the tensor de-composition of the vanilla Fully-connected layer, facilitates low-cost rescaling and global correlation modeling. BSM merges high-level features with granular ones adaptively. Furthermore, we provide an analysis based on the Univer-sal Approximation Theory of the NeRV system and vali-date the effectiveness of the proposed PNeRV. We conducted comprehensive experiments to demonstrate that PNeRV sur-passes the performance of contemporary NeRV models, achieving the best results in video regression on UVG and DAVIS under various metrics (PSNR, SSIM, LPIPS, and FVD). Compared to vanilla N eRV, P N eRV achieves <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> dB gain in PSNR and a 231% increase in FVD on UVG, along with <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> dB PSNR and 634% FVD increase on DAVIS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Tree-NeRV: Efficient Non-Uniform Sampling for Neural Video Representation via Tree-Structured Feature GridsJiancheng Zhao, Yifan Zhan, Qingtian Zhu, Mingze Ma 等ICCV 2025 · 被引用 4 次
- Exploring State-Space Models for Data-Specific Neural RepresentationsJinsung Lee, Suha KwakICLR 2026
- On Quantizing Neural Representation for Variable-Rate Video CodingJunqi Shi, Zhujia Chen, Hanfei Li, Qi Zhao 等ICLR 2025
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- HNeRV: A Hybrid Neural Representation for VideosHao Chen, Matthew Gwilliam, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2023
- DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic CodesHao Yan, Zhihui Ke, Xiaobo Zhou, Tie Qiu 等CVPR 2024 · 被引用 18 次
- DNeRV: Modeling Inherent Dynamics via Difference Neural Representation for VideosQi Zhao, M. Salman Asif, Zhan MaCVPR 2023
- MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal GuidanceJialong Guo, Ke Liu, Jiangchao Yao, Zhihua Wang 等AAAI 2025 · 被引用 7 次
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren 等NeurIPS 2021 · 被引用 430 次
