GIViC: Generative Implicit Video Compression
Ge Gao, Siyue Teng, Tianhao Peng, Fan Zhang, David Bull
Abstract
While video compression based on implicit neural representations (INRs) has recently demonstrated great potential, existing INR-based video codecs still cannot achieve state-of-the-art (SOTA) performance compared to their conventional or autoencoder-based counterparts given the same coding configuration. In this context, we propose a Generative Implicit Video Compression framework, GIViC, aiming at advancing the performance limits of this type of coding methods. GIViC draws inspiration from the remarkable ability of large language and diffusion models to capture long-range dependencies, a characteristic also inherent to Implicit Neural Representations (INRs). Through the newly designed implicit diffusion process, GIViC performs diffusive sampling across coarse-to-fine spatiotemporal decompositions, gradually progressing from coarsergrained full-sequence diffusion to finer-grained per-token diffusion. A novel Hierarchical Gated Linear Attention-based transformer (HGLA), is also integrated into the framework, which dual-factorizes global dependency modeling along scale and sequential axes. The proposed GIViC model has been benchmarked against SOTA conventional and neural codecs using a Random Access (RA) configuration (YUV 4:2:0, GOPSize=32), and yields BD-rate savings of and 8.52 % over VVC VTM, DCVC-FM and NVRC, respectively, on the UVG test set. As far as we are aware, GIViC is the first INR-based video codec that outperforms VTM, in terms of coding performance, based on the RA coding configuration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27ac0665-341e-4c15-bd8b-32f2a5533b0aCited by top-tier papers3
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- Real-Time Neural Video Compression with Unified Intra and Inter CodingHui Xiang, Yifan Bian, Li Li, Jingran Wu et al.CVPR 2026 · 5 citations
- Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation ModelsZongyu Guo, Jiajun He, Zhaoyang Jia, Xiaoyi Zhang et al.ICML 2026 · 1 citation
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
Related papers
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2024 · 44 citations
- Generative Neural Video Compression via Video Diffusion PriorQi Mao, Hao Cheng, Tinghan Yang, Libiao Jin et al.CVPR 2026 · 18 citations
- HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2023 · 132 citations
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 20 citations
- NeRV-Diffusion: Diffuse Implicit Neural Representation for Video SynthesisYixuan Ren, Hanyu Wang, Bo He, Hao Chen et al.ICLR 2026 · 1 citation
