Cross Modal Compression: Towards Human-comprehensible Semantic Compression
Jiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma, Wen Gao
摘要
Traditional image/video compression aims to reduce the transmission/storage cost with signal fidelity as high as possible. However, with the increasing demand for machine analysis and semantic monitoring in recent years, semantic fidelity rather than signal fidelity is becoming another emerging concern in image/video compression. With the recent advances in cross modal translation and generation, in this paper, we propose the cross modal compression (CMC), a semantic compression framework for visual data, to transform the high redundant visual data (such as image, video, etc.) into a compact, human-comprehensible domain (such as text, sketch, semantic map, attributions, etc.), while preserving the semantic. Specifically, we first formulate the CMC problem as a rate-distortion optimization problem. Secondly, we investigate the relationship with the traditional image/video compression and the recent feature compression frameworks, showing the difference between our CMC and these prior frameworks. Then we propose a novel paradigm for CMC to demonstrate its effectiveness. The qualitative and quantitative results show that our proposed CMC can achieve encouraging reconstructed results with an ultrahigh compression ratio, showing better compression performance than the widely used JPEG baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low BitratesZhitao Wang, Hengyu Man, Wenrui Li, Xingtao Wang 等AAAI 2026 · 被引用 5 次
- Blitzcrank: Fast Semantic Compression for In-memory Online Transaction ProcessingYiming Qiao, Yihan Gao, Huanchen ZhangVLDB 2024 · 被引用 3 次
相关 Paper
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative PriorRuoyu Feng, Yunpeng Qi, Jinming Liu, Yixin Gao 等NeurIPS 2025 · 被引用 5 次
- Multi-Modality Deep Network for Extreme Learned Image CompressionXuhao Jiang, Weimin Tan, Tian Tan, Bo Yan 等AAAI 2023 · 被引用 27 次
- Visual Redundancy Removal of Composite Images via Multimodal LearningWuyuan Xie, Shukang Wang, Rong Zhang, Miaohui WangACM MM 2023 · 被引用 1 次
- Semantic Scalable Image Compression with Cross-Layer PriorsHanyue Tu, Li Li, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 16 次
- Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionYuan Tian, Guo Lu, Guangtao Zhai, Zhiyong GaoICCV 2023 · 被引用 29 次
