MIMT: Masked Image Modeling Transformer for Video Compression
Jinxi Xiang, Kuan Tian, Jun Zhang
Abstract
Deep learning video compression outperforms its hand-craft counterparts with enhanced flexibility and capacity. One key component of the learned video codec is the autoregressive entropy model conditioned on spatial and temporal priors. Operating autoregressive on raster scanning order naively treats the context as unidirectional. This is neither efficient nor optimal, considering that conditional information probably locates at the end of the sequence. We thus introduce an entropy model based on a masked image modeling transformer (MIMT) to learn the spatial-temporal dependencies. Video frames are first encoded into sequences of tokens and then processed with the transformer encoder as priors. The transformer decoder learns the probability mass functions (PMFs) conditioned on the priors and masked inputs. Then it is capable of selecting optimal decoding orders without a fixed direction. During training, MIMT aims to predict the PMFs of randomly masked tokens by attending to tokens in all directions. This allows MIMT to capture the temporal dependencies from encoded priors and the spatial dependencies from the unmasked tokens, i.e., decoded tokens. At inference time, the model begins with generating PMFs of all masked tokens in parallel and then decodes the frame iteratively from the previously-selected decoded tokens (i.e., with high confidence). In addition, we improve the overall performance with more techniques, e.g., manifold conditional priors accumulating a long range of information, shifted window attention to reduce complexity. Extensive experiments demonstrate the proposed MIMT framework equipped with the new transformer entropy model achieves state-of-the-art performance on HEVC, UVG, and MCL-JCV datasets, generally outperforming the VVC in terms of PSNR and SSIM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get de6a5c7f-f5bb-44bf-a892-4eceadb6f7a5Cited by top-tier papers16
- HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2023 · 132 citations
- TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine PerceptionYi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao, Cheng Chien et al.ICCV 2023 · 59 citations
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2024 · 44 citations
- Neural Rate Control for Learned Video CompressionYiwei Zhang, Guo Lu, Yunuo Chen, Shen Wang et al.ICLR 2024 · 23 citations
- M2T: Masking Transformers Twice for Faster DecodingFabian Mentzer, Eirikur Agustsson, Michael TschannenICCV 2023 · 22 citations
Related papers
- Context Guided Transformer Entropy Modeling for Video CompressionJunlong Tong, Wei Zhang, Yaohui Jin, Xiaoyu ShenICCV 2025
- GIViC: Generative Implicit Video CompressionGe Gao, Siyue Teng, Tianhao Peng, Fan Zhang et al.ICCV 2025 · 4 citations
- Entroformer: A Transformer-based Entropy Model for Learned Image CompressionYichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan et al.ICLR 2022 · 194 citations
- BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangACM MM 2025 · 3 citations
- MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy CodingBowen Liu, Yu Chen, Rakesh Chowdary Machineni, Shiyu Liu et al.CVPR 2023
