Towards Unified Human Perception and Machine Understanding: Token Flow Guided Compression Framework
Li Xu, Yingfu Zhang, Kepeng Xu, Gang He, Yunsong Li
Abstract
With the rapid rise of Large Vision Language Models (LVLMs) for image understanding, the objective of image compression is gradually shifting from human visual perception to machine-oriented semantic understanding. However, conventional learned compression techniques are optimized for pixel-level fidelity and typically operate at fixed or rigid bitrate points, misaligned with the semantic consistency and flexible bitrate control. This gap becomes critical in ultra-low-bitrate regimes, where latent representations often ignore semantic relevance and struggle to disentangle meaningful content from redundant visual details as the bitrate varies. To address these challenges, we develop a token-based flexible compression framework, Token Flow Guided Compression (TFGC), which unifies humanand machine-oriented objectives. TFGC supports variable bitrate control in ultra-low bitrate regimes and enables LVLMs to directly process compressed tokens without image reconstruction. Specifically, we explore the token flow phenomenon in 1D token sequences and exploit it to design token flow propagation, which predicts missing tokens by propagating contextual information from unmasked tokens. Moreover, token semantic guidance aligns compressed representations with LVLM semantic space, while a progressive semantic alignment training strategy further bridges the gap between perceptual reconstruction and semantic reasoning. Experiments show that our framework achieves state-of-the-art LVLMs understanding at comparable bitrates while maintaining satisfactory perceptual quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04dbb51c-4aee-4ba4-a641-0bf207427439Builds on26
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
Related papers
- CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMsJingyu Lei, Gaoang Wang, Der-Horng LeeCVPR 2026 · 1 citation
- PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language ModelsChenyu Yang, Xuan Dong, Xizhou Zhu, Weijie Su et al.CVPR 2025
- HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image CompressionLei Lu, Yanyue Xie, Wei Jiang, Wei Wang et al.ACM MM 2024 · 9 citations
- T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low BitratesZhitao Wang, Hengyu Man, Wenrui Li, Xingtao Wang et al.AAAI 2026 · 5 citations
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video UnderstandingZheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao et al.NeurIPS 2025 · 5 citations
