Cross Modal Compression: Towards Human-comprehensible Semantic Compression
Jiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma, Wen Gao
Abstract
Traditional image/video compression aims to reduce the transmission/storage cost with signal fidelity as high as possible. However, with the increasing demand for machine analysis and semantic monitoring in recent years, semantic fidelity rather than signal fidelity is becoming another emerging concern in image/video compression. With the recent advances in cross modal translation and generation, in this paper, we propose the cross modal compression (CMC), a semantic compression framework for visual data, to transform the high redundant visual data (such as image, video, etc.) into a compact, human-comprehensible domain (such as text, sketch, semantic map, attributions, etc.), while preserving the semantic. Specifically, we first formulate the CMC problem as a rate-distortion optimization problem. Secondly, we investigate the relationship with the traditional image/video compression and the recent feature compression frameworks, showing the difference between our CMC and these prior frameworks. Then we propose a novel paradigm for CMC to demonstrate its effectiveness. The qualitative and quantitative results show that our proposed CMC can achieve encouraging reconstructed results with an ultrahigh compression ratio, showing better compression performance than the widely used JPEG baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low BitratesZhitao Wang, Hengyu Man, Wenrui Li, Xingtao Wang et al.AAAI 2026 · 5 citations
- Blitzcrank: Fast Semantic Compression for In-memory Online Transaction ProcessingYiming Qiao, Yihan Gao, Huanchen ZhangVLDB 2024 · 3 citations
Related papers
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative PriorRuoyu Feng, Yunpeng Qi, Jinming Liu, Yixin Gao et al.NeurIPS 2025 · 5 citations
- Multi-Modality Deep Network for Extreme Learned Image CompressionXuhao Jiang, Weimin Tan, Tian Tan, Bo Yan et al.AAAI 2023 · 27 citations
- Visual Redundancy Removal of Composite Images via Multimodal LearningWuyuan Xie, Shukang Wang, Rong Zhang, Miaohui WangACM MM 2023 · 1 citation
- Semantic Scalable Image Compression with Cross-Layer PriorsHanyue Tu, Li Li, Wengang Zhou, Houqiang LiACM MM 2021 · 16 citations
- Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionYuan Tian, Guo Lu, Guangtao Zhai, Zhiyong GaoICCV 2023 · 29 citations
