Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
Qi Yang, Chenghao Zhang, Lubin Fan, Kun Ding, Jieping Ye, Shiming Xiang
摘要
Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented Generation (RAG). However, existing methods still face challenges, such as the scarcity of knowledge containing reasoning examples and erratic responses from retrieved knowledge. To address these issues, in this study, we propose a multimodal RAG framework, termed RCTS, which enhances LVLMs by constructing a Reasoning Context-enriched knowledge base and a Tree Search re-ranking method. Specifically, we introduce a self-consistent evaluation mechanism to enrich the knowledge base with intrinsic reasoning patterns. We further propose a Monte Carlo Tree Search with Heuristic Rewards (MCTS-HR) to prioritize the most relevant examples. This ensures that LVLMs can leverage high-quality contextual reasoning for better and more consistent responses. Extensive experiments demonstrate that our framework achieves state-of-the-art performance across multiple VQA datasets, significantly outperforming both In-Context Learning (ICL) and Vanilla-RAG methods. It highlights the effectiveness of our knowledge base and reranking method in improving LVLMs. Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger Similarity Search Top-N Contexts Embeddings MCTS Re-ranking Knowledge Base Q: Which rhetorical appeal is primarily used in this ad? A: logos (reason) C: Identify the type of appeal: The ad claims that the Vilaplus vacuum picks up more dirt than User Question In this food web, which organism contains matter that eventually moves to the bat star? Below is a food web from an ocean ecosystem in Monterey Bay. LVLM Reasoning Contexts Answer User Question Retrieval Model MLP Retrieval Model #3 #1 Q: Based on the arrows, Below is a food web from an ocean ecosystem. A: kelp bass C: Identify the type of question and context: The user is asked to identif y which living ... Q: Which of these organisms contains matter that was once part of the bear sedge? A: snowy owl C: Identify the question: The question asks which organism contains matter that … 1. Identify the Image Content: The image shows the Great Wall of China, which is a series of fortifications built over several centuries, primarily during the Qin and Han dynasties. 2. Understand the Historical Context: The Great Wall was ... Reasoning Context 1. Identify the Image Content: The image shows the Great Wall of China, which is a series of fortifications built over several centuries, primarily during the Qin and Han dynasties. 2. Understand the Historical Context: The Great Wall was ... Reasoning Context 1. Understand the Image: The image depicts the Great Wall of China, a historical structure that stretches across the mountainous terrain. 2. Contextual Knowledge Integration: The Great Wall was constructed and extended over ...
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic BridgingMing Zhong, Yuanlei Wang, Liuzhou Zhang, Ruichuan An 等CVPR 2026 · 被引用 2 次
- Dual-Latent Memory Routing for Vision-Language ReasoningHao-Xuan Ma, Jin-Fei Qi, YiCheng Xiao, Han-Jia YeICML 2026
它引用的顶会 Paper14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 被引用 234 次
相关 Paper
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi 等CVPR 2026 · 被引用 11 次
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li 等AAAI 2026 · 被引用 1 次
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen 等NeurIPS 2025 · 被引用 76 次
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward MechanismCongzhi Zhang, Jiawei Peng, Zhenglin Wang, Yilong Lai 等ACL 2025 · 被引用 6 次
- SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet KnowledgeChuanhao Li, Zhen Li, Chenchen Jing, Shuo Liu 等NeurIPS 2024 · 被引用 26 次
