Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye
摘要
Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-augmented generation (RAG) achieves significant advances in this task by combining knowledge-base querying, it still struggles with the quality of multimodal queries and the relevance of retrieved results. To overcome these challenges, we propose a novel three-stage method, termed Wiki-PRF, including Processing, Retrieval and Filtering stages. The processing stage dynamically invokes visual tools to extract precise multimodal information for retrieval. The retrieval stage integrates visual and text features to achieve multimodal knowledge retrieval. The filtering stage performs relevance filtering and concentration on retrieval results. To this end, we introduce a visual language model trained with answer accuracy and format consistency as reward signals via a reinforcement learning manner. This enhances the model's reasoning, tool invocation for accurate queries, and filtering of irrelevant content. Experiments on benchmark datasets Question: Which animals eats the fruits / leaves of this plant?
Vanilla RAG: The fruits and leaves of Monarda fistulosa//...//many insects and small birds may utilize the nectar and pollen from the flowers for food.
Based on the information provided in the context, the article does not specify which animals. //....// Additionally, herbivorous mammals are said to avoid the foliage due to its odor.
The passage does not provide//...//The known pollinators for this plant are bees, butterflies, and hummingbirds, as well as DOUFIOREA MONARDAE, which itself is a pollinator.
Question: How many meters high can this plant grow?
Vanilla RAG: The image shows Fuchsia magellanica, which typically grows to heights of 1 to 2 meters.
This plant can grow up to 10 meters in height in frost-free climates.
The plant Fuchsia magellanica can grow up to 3.0 meters high in frost-free climates and up to 1.5 meters high in colder conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi 等CVPR 2026 · 被引用 11 次
- CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question AnsweringYuyang Hong, Jiaqi Gu, Yujing Lou, Lubin Fan 等CVPR 2026 · 被引用 2 次
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee 等CVPR 2026 · 被引用 2 次
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao 等ICML 2026 · 被引用 1 次
- Beyond Single-View Indexing: Structure-Aware Multi-View Retrieval for Knowledge-Based VQAHao Wang, Xujia Li, Lei ChenICML 2026
它引用的顶会 Paper22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models StrongerQi Yang, Chenghao Zhang, Lubin Fan, Kun Ding 等ICML 2025
- Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model ReasoningMingyuan Wu, Jize Jiang, Haozhen Zheng, Meitang Li 等EMNLP 2025
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement LearningLu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu 等NeurIPS 2025 · 被引用 4 次
- KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree SearchHaoran Luo, Haihong E, Yikai Guo, Qika Lin 等ICML 2025 · 被引用 1 次
- CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question AnsweringHao Yang, Zhiyu Yang, Xupeng Zhang, Wei Wei 等WWW 2026
