FLAVC: Learned Video Compression with Feature Level Attention
Chun Zhang, Heming Sun, Jiro Katto
Abstract
Learned Video Compression (LVC) aims to reduce redundancy in sequential data through deep learning approaches. Recent advances have significantly boosted LVC performance by shifting compression operations to the feature domain, often combining Motion Estimation and Motion Compensation modules (MEMC) with CNN-based context extraction. However, reliance on motions and convolutiondriven context models limits generalizability and global perception. To address these issues, we propose a Featurelevel Attention (FLA) module within a Transformer-based framework that explicitly perceives full-frame, thus bypassing confined motion signatures. FLA accomplishes global perception by converting high-level local patch embeddings into one-dimensional batch-wise vectors and replacing traditional attention weights to a global context matrix. Additionally, a dense overlapping patcher (DP) is introduced to retain local features before embedding projection. Furthermore, a Transformer-CNN mixed encoder is applied to alleviate the spatial feature bottleneck without increasing latent size. Experiments demonstrate excellent generalizability with universally efficient redundancy reduction in different scenarios. Extensive tests on four video compression datasets show that our method achieves state-of-theart Rate-Distortion performance compared to existing LVC methods and traditional codecs. A down-scaled version of our model reduced computation overhead by a great margin while maintained strong performance. The code is available at https://github.com/Z-CV-code/FLAVC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb2420e7-cd7f-4b99-a1aa-e8be63f88ac2Cited by top-tier papers1
Ask how each one uses itBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
Related papers
- FVC: A New Framework Towards Deep Video Compression in Feature SpaceZhihao Hu, Guo Lu, Dong XuCVPR 2021
- BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangACM MM 2025 · 3 citations
- Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersZhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang et al.ACM MM 2023 · 20 citations
- ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangCVPR 2025
- MoVie: Multimodal Video Compression with Text GuidanceJiaqi Hu, Haoji Hu, Heming Sun, Lianrui MuICML 2026
