PerceptionGPT: Effectively Fusing Visual Perception Into LLM
Renjie Pi, Lewei Yao, Jiahui Gao, Jipeng Zhang, Tong Zhang
摘要
The integration of visual inputs with large language models (LLMs) has led to remarkable advancements in multi-modal capabilities, giving rise to vision large language models (VLLMs). However, effectively harnessing LLMs for intricate visual perception tasks, such as detection and segmentation, remains a challenge. Conventional approaches achieve this by transforming perception signals (e.g., bounding boxes, segmentation masks) into sequences of discrete tokens, which struggle with the precision errors and introduces further complexities for training. In this paper, we present a novel end-to-end framework named Per-ceptionGPT, which represent the perception signals using LLM's dynamic token embedding. Specifically, we leverage lightweight encoders and decoders to handle the perception signals in LLM's embedding space, which takes advantage of the representation power of the high-dimensional token embeddings. Our approach significantly eases the training difficulties associated with the discrete representations in prior methods. Furthermore, owing to our compact representation, the inference speed is also greatly boosted. Consequently, PerceptionGPT enables accurate, flexible and efficient handling of complex perception signals. We validate the effectiveness of our approach through extensive experiments. The results demonstrate significant improvements over previous methods with only 4% trainable parameters and less than 25% training time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu 等NeurIPS 2025 · 被引用 33 次
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language InterfaceHao Tang, Chen-Wei Xie, Haiyang Wang, Xiaoyi Bao 等NeurIPS 2025 · 被引用 30 次
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning SegmentationZhenyu Lu, Liupeng Li, Jinpeng Wang, Yan Feng 等ICLR 2026 · 被引用 10 次
- KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionJie Yang, Wang Zeng, Sheng Jin, Lumin Xu 等NeurIPS 2024 · 被引用 9 次
- Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language ModelsCong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng 等ICCV 2025 · 被引用 8 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMsYongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu 等ICLR 2026 · 被引用 22 次
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei 等ICCV 2025 · 被引用 1 次
- HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual PerceiverCong Wei, Yujie Zhong, Haoxian Tan, Yong Liu 等CVPR 2025
- DenseMLLM: Standard Multimodal LLMs for Dense PredictionYi Li, Hongze Shen, Lexiang Tang, Xin Li 等ICML 2026
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang 等NeurIPS 2024 · 被引用 24 次
