FLAVC: Learned Video Compression with Feature Level Attention
Chun Zhang, Heming Sun, Jiro Katto
摘要
Learned Video Compression (LVC) aims to reduce redundancy in sequential data through deep learning approaches. Recent advances have significantly boosted LVC performance by shifting compression operations to the feature domain, often combining Motion Estimation and Motion Compensation modules (MEMC) with CNN-based context extraction. However, reliance on motions and convolutiondriven context models limits generalizability and global perception. To address these issues, we propose a Featurelevel Attention (FLA) module within a Transformer-based framework that explicitly perceives full-frame, thus bypassing confined motion signatures. FLA accomplishes global perception by converting high-level local patch embeddings into one-dimensional batch-wise vectors and replacing traditional attention weights to a global context matrix. Additionally, a dense overlapping patcher (DP) is introduced to retain local features before embedding projection. Furthermore, a Transformer-CNN mixed encoder is applied to alleviate the spatial feature bottleneck without increasing latent size. Experiments demonstrate excellent generalizability with universally efficient redundancy reduction in different scenarios. Extensive tests on four video compression datasets show that our method achieves state-of-theart Rate-Distortion performance compared to existing LVC methods and traditional codecs. A down-scaled version of our model reduced computation overhead by a great margin while maintained strong performance. The code is available at https://github.com/Z-CV-code/FLAVC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma 等CVPR 2022 · 被引用 363 次
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 被引用 202 次
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles 等NeurIPS 2022 · 被引用 155 次
相关 Paper
- FVC: A New Framework Towards Deep Video Compression in Feature SpaceZhihao Hu, Guo Lu, Dong XuCVPR 2021
- BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangACM MM 2025 · 被引用 3 次
- Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersZhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang 等ACM MM 2023 · 被引用 20 次
- ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangCVPR 2025
- MoVie: Multimodal Video Compression with Text GuidanceJiaqi Hu, Haoji Hu, Heming Sun, Lianrui MuICML 2026
