All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path Aggregation
Xu Zhang, Peiyao Guo, Ming Lu, Zhan Ma
摘要
Image coding for multi-task applications, catering to both human perception and machine vision, has been extensively investigated. Existing methods often rely on multiple task-specific encoder-decoder pairs, leading to high overhead of parameter and bitrate usage, or face challenges in multi-objective optimization under a unified representation, failing to achieve both performance and efficiency. To this end, we propose Multi-Path Aggregation (MPA) integrated into existing coding models for joint human-machine vision, unifying the feature representation with an all-in-one architecture. MPA employs a predictor to allocate latent features among task-specific paths based on feature importance varied across tasks, maximizing the utility of shared features while preserving task-specific features for subsequent refinement. Leveraging feature correlations, we develop a two-stage optimization strategy to alleviate multi-task performance degradation. Upon the reuse of shared features, as low as 1.89% parameters are further augmented and fine-tuned for a specific task, which completely avoids extensive optimization of the entire model. Experimental results show that MPA achieves performance comparable to state-of-the-art methods in both task-specific and multi-objective optimization across human viewing and machine analysis tasks. Moreover, our all-in-one design supports seamless transitions between human- and machine-oriented reconstruction, enabling task-controllable interpretation without altering the unified model. Code is available at https://github.com/NJUVISION/MPA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkChangsheng Gao, Yifan Ma, Qiaoxi Chen, Yenan Xu 等ICCV 2025 · 被引用 3 次
- Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation CompressionZhimeng Huang, Rongao Yuan, Junlong Gao, Qi Mao 等CVPR 2026
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
相关 Paper
- ICMH-Net: Neural Image Compression Towards both Machine Vision and Human VisionLei Liu, Zhihao Hu, Zhenghao Chen, Dong XuACM MM 2023 · 被引用 21 次
- ALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary LatencyDong Wei, Xiaoning Sun, Xizhan Gao, Shengxiang Hu 等CVPR 2025
- Lossy Common Information in a Learnable Gray-Wyner NetworkAnderson de Andrade, Alon Harell, Ivan V. BajicICLR 2026 · 被引用 1 次
- Learning Versatile Neural Architectures by Propagating Network CodesMingyu Ding, Yuqi Huo, Haoyu Lu, Linjie Yang 等ICLR 2022 · 被引用 14 次
- Transfer Vision Patterns for Multi-Task Pixel LearningXiaoya Zhang, Ling Zhou, Yong Li, Zhen Cui 等ACM MM 2021 · 被引用 11 次
