ICMH-Net: Neural Image Compression Towards both Machine Vision and Human Vision
Lei Liu, Zhihao Hu, Zhenghao Chen, Dong Xu
摘要
Neural image compression has gained significant attention thanks to the remarkable success of deep neural networks. However, most existing neural image codecs focus solely on improving human vision perception. In this work, our objective is to enhance image compression methods for both human vision quality and machine vision tasks simultaneously. To achieve this, we introduce a novel approach to Partition, Transmit, Reconstruct, and Aggregate (PTRA) the latent representation of images to balance the optimizations for both aspects. By employing our method as a module in existing neural image codecs, we create a latent representation predictor that dynamically manages the bit-rate cost for machine vision tasks. To further improve the performance of auto-regressive-based coding techniques, we enhance our hyperprior network and predictor module with context modules, resulting in a reduction in bit-rate. The extensive experiments conducted on various machine vision benchmarks such as ILSVRC 2012, VOC 2007, VOC 2012, and COCO demonstrate the superiority of our newly proposed image compression framework. It outperforms existing neural image compression methods in multiple machine vision tasks including classification, segmentation, and detection, while maintaining high-quality image reconstruction for human vision.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersZhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang 等ACM MM 2023 · 被引用 20 次
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video CompressionZhenghao Chen, Luping Zhou, Zhihao Hu, Dong XuACM MM 2024 · 被引用 14 次
- Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP SupervisionKangsheng Yin, Quan Liu, Xuelin Shen, Yulin He 等AAAI 2025 · 被引用 6 次
- 3D Gaussian Splatting Data Compression with Mixture of PriorsLei Liu, Zhenghao Chen, Dong XuACM MM 2025 · 被引用 4 次
- Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image CompressionShiyin Jiang, Wei Long, Minghao Han, Zhenghao Chen 等CVPR 2026 · 被引用 3 次
相关 Paper
- All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path AggregationXu Zhang, Peiyao Guo, Ming Lu, Zhan MaNeurIPS 2024 · 被引用 20 次
- Semantic Scalable Image Compression with Cross-Layer PriorsHanyue Tu, Li Li, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 16 次
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 被引用 143 次
- Bridging Compressed Image Latents and Multimodal Large Language ModelsChia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen 等ICLR 2025
- Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionChuqin Zhou, Guo Lu, Jiangchuan Li, Xiangyu Chen 等AAAI 2025 · 被引用 3 次
