PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement
Zhanfeng Feng, Long Peng, Xin Di, Yong Guo, Wenbo Li, Yulun Zhang, Renjing Pei, Yang Wang, Yang Cao, Zheng-Jun Zha
Abstract
Multi-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhancement methods have achieved impressive performance, their computational and memory demands hinder deployment on edge devices. Quantization offers a practical solution by reducing the bit-width of weights and activations to improve efficiency. However, directly applying existing quantization methods to video enhancement tasks often leads to significant performance degradation and loss of fine details. This stems from two limitations: (a) inability to allocate varying representational capacity across frames, which results in suboptimal dynamic range adaptation; (b) over-reliance on full-precision teachers, which limits the learning of low-bit student models. To tackle these challenges, we propose a novel quantization method for video enhancement: Progressive Multi-Frame Quantization for Video Enhancement (PMQ-VE). This framework features a coarse-to-fine two-stage process: Backtracking-based Multi-Frame Quantization (BMFQ) and Progressive Multi-Teacher Distillation (PMTD). BMFQ utilizes a percentile-based initialization and iterative search with pruning and backtracking for robust clipping bounds. PMTD employs a progressive distillation strategy with both full-precision and multiple high-bit (INT) teachers to enhance low-bit models' capacity and quality. Extensive experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance across multiple tasks and benchmarks.The code will be made publicly available at: https://github.com/xiaoBIGfeng/PMQ-VE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07a7f3bd-b966-4a47-bd1d-97e03af087beCited by top-tier papers2
- Depth-Synergized Mamba Meets Memory Experts for All-Day Image Reflection SeparationSiyan Fang, Long Peng, Yuntao Wang, Ruonan Wei et al.AAAI 2026 · 5 citations
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi et al.CVPR 2026
Builds on38
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Channel Attention Is All You Need for Video Frame InterpolationMyungsub Choi, Heewon Kim, Bohyung Han, Ning Xu et al.AAAI 2020 · 362 citations
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 172 citations
- Rethinking Alignment in Video Super-Resolution TransformersShuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang et al.NeurIPS 2022 · 134 citations
Related papers
- Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersWeilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li et al.ICML 2025
- Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality AssessmentKun Yuan, Zishang Kong, Chuanchuan Zheng, Ming Sun et al.ACM MM 2023 · 15 citations
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo et al.NeurIPS 2022 · 17 citations
- QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention SparsificationWeilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu et al.ICLR 2026 · 8 citations
