Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye
Abstract
Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-augmented generation (RAG) achieves significant advances in this task by combining knowledge-base querying, it still struggles with the quality of multimodal queries and the relevance of retrieved results. To overcome these challenges, we propose a novel three-stage method, termed Wiki-PRF, including Processing, Retrieval and Filtering stages. The processing stage dynamically invokes visual tools to extract precise multimodal information for retrieval. The retrieval stage integrates visual and text features to achieve multimodal knowledge retrieval. The filtering stage performs relevance filtering and concentration on retrieval results. To this end, we introduce a visual language model trained with answer accuracy and format consistency as reward signals via a reinforcement learning manner. This enhances the model's reasoning, tool invocation for accurate queries, and filtering of irrelevant content. Experiments on benchmark datasets Question: Which animals eats the fruits / leaves of this plant?
Vanilla RAG: The fruits and leaves of Monarda fistulosa//...//many insects and small birds may utilize the nectar and pollen from the flowers for food.
Based on the information provided in the context, the article does not specify which animals. //....// Additionally, herbivorous mammals are said to avoid the foliage due to its odor.
The passage does not provide//...//The known pollinators for this plant are bees, butterflies, and hummingbirds, as well as DOUFIOREA MONARDAE, which itself is a pollinator.
Question: How many meters high can this plant grow?
Vanilla RAG: The image shows Fuchsia magellanica, which typically grows to heights of 1 to 2 meters.
This plant can grow up to 10 meters in height in frost-free climates.
The plant Fuchsia magellanica can grow up to 3.0 meters high in frost-free climates and up to 1.5 meters high in colder conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58c3e83a-bd22-4405-8473-d227635e92caCited by top-tier papers5
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi et al.CVPR 2026 · 11 citations
- CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question AnsweringYuyang Hong, Jiaqi Gu, Yujing Lou, Lubin Fan et al.CVPR 2026 · 2 citations
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee et al.CVPR 2026 · 2 citations
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao et al.ICML 2026 · 1 citation
- Beyond Single-View Indexing: Structure-Aware Multi-View Retrieval for Knowledge-Based VQAHao Wang, Xujia Li, Lei ChenICML 2026
Builds on22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models StrongerQi Yang, Chenghao Zhang, Lubin Fan, Kun Ding et al.ICML 2025
- Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model ReasoningMingyuan Wu, Jize Jiang, Haozhen Zheng, Meitang Li et al.EMNLP 2025
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement LearningLu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu et al.NeurIPS 2025 · 4 citations
- KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree SearchHaoran Luo, Haihong E, Yikai Guo, Qika Lin et al.ICML 2025 · 1 citation
- CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question AnsweringHao Yang, Zhiyu Yang, Xupeng Zhang, Wei Wei et al.WWW 2026
