CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
Junyan Li, Delin Chen, Yining Hong, Zhenfang Chen, Peihao Chen, Yikang Shen, Chuang Gan
Abstract
A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large visionlanguage foundation models (VLMs) fall short of such compositional abilities due to their "bag-of-words" behaviors and inability to construct words that correctly represent visual entities and the relations among the entities. To this end, we propose CoVLM, which can guide the LLM to explicitly compose visual entities and relationships among the text and dynamically communicate with the vision encoder and detection network to achieve vision-language communicative decoding. Specifically, we first devise a set of novel communication tokens for the LLM, for dynamic communication between the visual detection system and the language system. A communication token is generated by the LLM following a visual entity or a relation, to inform the detection network to propose regions that are relevant to the sentence generated so far. The proposed regions-of-interests (ROIs) are then fed back into the LLM for better language generation contingent on the relevant regions. The LLM is thus able to compose the visual entities and relationships through the communication tokens. The vision-to-language and language-to-vision communication are iteratively performed until the entire sentence is generated. Our framework seamlessly bridges the gap between visual perception and LLMs and outperforms previous VLMs by a large margin on compositional reasoning benchmarks (e.g., ∼ 20% in HICO-DET mAP, ∼ 14% in Cola top-1 accuracy, and ∼ 3% on ARO top-1 accuracy). We also achieve competitive performances on traditional vision-language tasks such as referring expression comprehension and visual question answering 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cac8b48-0d5b-4672-9e8b-064cab95ca5aCited by top-tier papers10
- Kestrel: 3D Multimodal LLM for Part-Aware Grounded DescriptionMahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr et al.ICCV 2025 · 9 citations
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo et al.ICCV 2025 · 8 citations
- Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorJiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei et al.ACM MM 2024 · 3 citations
- Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language ModelsQuang-Hung Le, Long Hoang Dang, Ngan Hoang Le, Truyen Tran et al.AAAI 2025 · 3 citations
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingTa Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi et al.ICCV 2025 · 2 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Enhancing Advanced Visual Reasoning Ability of Large Language ModelsZhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang et al.EMNLP 2024 · 10 citations
- COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision-Language ModelsSanchit Sinha, Guangzhi Xiong, Aidong ZhangEMNLP 2025 · 1 citation
- Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional UnderstandingLe Zhang, Rabiul Awal, Aishwarya AgrawalCVPR 2024 · 7 citations
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
- Causal Graphical Models for Vision-Language Compositional UnderstandingFiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi et al.ICLR 2025
