Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Zihan Zheng, Yuan Zhang, Yan Lu
Abstract
While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still suffer from artifacts due to limited generation capacity, whereas others leverage pretrained diffusion models to improve quality at the cost of high sampling complexity. To overcome these challenges, we propose S 2 VC, a Single-Step diffusion-based Video Codec that integrates a conditional coding framework with an efficient singlestep diffusion generator, enabling realistic reconstruction at low bitrates with reduced sampling cost. Recognizing the importance of semantic conditioning in single-step diffusion, we introduce Contextual Semantic Guidance to extract frame-adaptive semantics from buffered features. This guidance replaces text captions with efficient, fine-grained conditioning, thereby improving generation realism. In addition, Temporal Consistency Guidance is incorporated into the diffusion U-Net to enforce temporal coherence across frames and ensure stable generation. Extensive experiments show that S 2 VC delivers state-of-the-art perceptual quality with an average bitrate saving of 51.62% over prior perceptual method, underscoring the promise of single-step diffusion for efficient, high-quality video compression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- Generative Video Compression with One-Dimensional Latent RepresentationZihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li et al.CVPR 2026 · 5 citations
Builds on25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image CompressionZheng Chen, Mingde Zhou, Jinpei Guo, Jiale Yuan et al.AAAI 2026 · 1 citation
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An et al.ICLR 2025
- Generative Neural Video Compression via Video Diffusion PriorQi Mao, Hao Cheng, Tinghan Yang, Libiao Jin et al.CVPR 2026 · 18 citations
- StableCodec: Taming One-Step Diffusion for Extreme Image CompressionTianyu Zhang, Xin Luo, Li Li, Dong LiuICCV 2025 · 7 citations
- PICD: Versatile Perceptual Image Compression with Diffusion RenderingTongda Xu, Jiahao Li, Bin Li, Yan Wang et al.CVPR 2025
