MoVie: Multimodal Video Compression with Text Guidance
Jiaqi Hu, Haoji Hu, Heming Sun, Lianrui Mu
Abstract
Most deep video codecs emphasize low-level motion modeling and remain largely semantics-agnostic, which can degrade perceptual quality in complex scenes. We propose MoVie, a Multimodal Video compression framework built on a Text-guided Video Transformer–CNN Mixed block (Text-VideoTCM). MoVie adopts a video-centric architecture that jointly models local spatial structures and temporal dynamics via window-based processing, delivering a favorable computation--perception trade-off. To incorporate semantics, we introduce dual-stage text fusion with Extractor and Injector modules. We further present history-conditioned coding that leverages both previous and aggregated historical frames, and a spatial--channel factorized entropy model that estimates probabilities over spatial neighborhoods and channel groups for adaptive bit allocation. Together, these designs reduce redundancy and improve rate control and temporal coherence, yielding reconstructions at low bitrates. On UVG and MCL-JCV, MoVie achieves 50.23% BD-rate for FID and 14.64% for LPIPS (VGGNet) relative to HM, while requiring only 55.76% of DCVC-FM's per-pixel kMACs. A human perceptual study further confirms consistent subjective preference over strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 423ccb1e-6c84-41cb-ac1a-92190e8980e0Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
Related papers
- MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy CodingBowen Liu, Yu Chen, Rakesh Chowdary Machineni, Shiyu Liu et al.CVPR 2023
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- MIMT: Masked Image Modeling Transformer for Video CompressionJinxi Xiang, Kuan Tian, Jun ZhangICLR 2023
- Context Guided Transformer Entropy Modeling for Video CompressionJunlong Tong, Wei Zhang, Yaohui Jin, Xiaoyu ShenICCV 2025
