ICMH-Net: Neural Image Compression Towards both Machine Vision and Human Vision
Lei Liu, Zhihao Hu, Zhenghao Chen, Dong Xu
Abstract
Neural image compression has gained significant attention thanks to the remarkable success of deep neural networks. However, most existing neural image codecs focus solely on improving human vision perception. In this work, our objective is to enhance image compression methods for both human vision quality and machine vision tasks simultaneously. To achieve this, we introduce a novel approach to Partition, Transmit, Reconstruct, and Aggregate (PTRA) the latent representation of images to balance the optimizations for both aspects. By employing our method as a module in existing neural image codecs, we create a latent representation predictor that dynamically manages the bit-rate cost for machine vision tasks. To further improve the performance of auto-regressive-based coding techniques, we enhance our hyperprior network and predictor module with context modules, resulting in a reduction in bit-rate. The extensive experiments conducted on various machine vision benchmarks such as ILSVRC 2012, VOC 2007, VOC 2012, and COCO demonstrate the superiority of our newly proposed image compression framework. It outperforms existing neural image compression methods in multiple machine vision tasks including classification, segmentation, and detection, while maintaining high-quality image reconstruction for human vision.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersZhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang et al.ACM MM 2023 · 20 citations
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video CompressionZhenghao Chen, Luping Zhou, Zhihao Hu, Dong XuACM MM 2024 · 14 citations
- Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP SupervisionKangsheng Yin, Quan Liu, Xuelin Shen, Yulin He et al.AAAI 2025 · 6 citations
- 3D Gaussian Splatting Data Compression with Mixture of PriorsLei Liu, Zhenghao Chen, Dong XuACM MM 2025 · 4 citations
- Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image CompressionShiyin Jiang, Wei Long, Minghao Han, Zhenghao Chen et al.CVPR 2026 · 3 citations
Related papers
- All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path AggregationXu Zhang, Peiyao Guo, Ming Lu, Zhan MaNeurIPS 2024 · 20 citations
- Semantic Scalable Image Compression with Cross-Layer PriorsHanyue Tu, Li Li, Wengang Zhou, Houqiang LiACM MM 2021 · 16 citations
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 143 citations
- Bridging Compressed Image Latents and Multimodal Large Language ModelsChia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen et al.ICLR 2025
- Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionChuqin Zhou, Guo Lu, Jiangchuan Li, Xiangyu Chen et al.AAAI 2025 · 3 citations
