Lune

ICML2026顶会

MoVie: Multimodal Video Compression with Text Guidance

Jiaqi Hu, Haoji Hu, Heming Sun, Lianrui Mu

出版方
2026年份

摘要

Most deep video codecs emphasize low-level motion modeling and remain largely semantics-agnostic, which can degrade perceptual quality in complex scenes. We propose MoVie, a Multimodal Video compression framework built on a Text-guided Video Transformer–CNN Mixed block (Text-VideoTCM). MoVie adopts a video-centric architecture that jointly models local spatial structures and temporal dynamics via window-based processing, delivering a favorable computation--perception trade-off. To incorporate semantics, we introduce dual-stage text fusion with Extractor and Injector modules. We further present history-conditioned coding that leverages both previous and aggregated historical frames, and a spatial--channel factorized entropy model that estimates probabilities over spatial neighborhoods and channel groups for adaptive bit allocation. Together, these designs reduce redundancy and improve rate control and temporal coherence, yielding reconstructions at low bitrates. On UVG and MCL-JCV, MoVie achieves −-50.23% BD-rate for FID and −-14.64% for LPIPS (VGGNet) relative to HM, while requiring only 55.76% of DCVC-FM's per-pixel kMACs. A human perceptual study further confirms consistent subjective preference over strong baselines.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖