All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path Aggregation
Xu Zhang, Peiyao Guo, Ming Lu, Zhan Ma
Abstract
Image coding for multi-task applications, catering to both human perception and machine vision, has been extensively investigated. Existing methods often rely on multiple task-specific encoder-decoder pairs, leading to high overhead of parameter and bitrate usage, or face challenges in multi-objective optimization under a unified representation, failing to achieve both performance and efficiency. To this end, we propose Multi-Path Aggregation (MPA) integrated into existing coding models for joint human-machine vision, unifying the feature representation with an all-in-one architecture. MPA employs a predictor to allocate latent features among task-specific paths based on feature importance varied across tasks, maximizing the utility of shared features while preserving task-specific features for subsequent refinement. Leveraging feature correlations, we develop a two-stage optimization strategy to alleviate multi-task performance degradation. Upon the reuse of shared features, as low as 1.89% parameters are further augmented and fine-tuned for a specific task, which completely avoids extensive optimization of the entire model. Experimental results show that MPA achieves performance comparable to state-of-the-art methods in both task-specific and multi-objective optimization across human viewing and machine analysis tasks. Moreover, our all-in-one design supports seamless transitions between human- and machine-oriented reconstruction, enabling task-controllable interpretation without altering the unified model. Code is available at https://github.com/NJUVISION/MPA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fb8c659-f107-440b-b58a-6ef33bf5acf4Cited by top-tier papers2
- Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkChangsheng Gao, Yifan Ma, Qiaoxi Chen, Yenan Xu et al.ICCV 2025 · 3 citations
- Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation CompressionZhimeng Huang, Rongao Yuan, Junlong Gao, Qi Mao et al.CVPR 2026
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- ICMH-Net: Neural Image Compression Towards both Machine Vision and Human VisionLei Liu, Zhihao Hu, Zhenghao Chen, Dong XuACM MM 2023 · 21 citations
- ALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary LatencyDong Wei, Xiaoning Sun, Xizhan Gao, Shengxiang Hu et al.CVPR 2025
- Lossy Common Information in a Learnable Gray-Wyner NetworkAnderson de Andrade, Alon Harell, Ivan V. BajicICLR 2026 · 1 citation
- Learning Versatile Neural Architectures by Propagating Network CodesMingyu Ding, Yuqi Huo, Haoyu Lu, Linjie Yang et al.ICLR 2022 · 14 citations
- Transfer Vision Patterns for Multi-Task Pixel LearningXiaoya Zhang, Ling Zhou, Yong Li, Zhen Cui et al.ACM MM 2021 · 11 citations
