Beyond Single-View Indexing: Structure-Aware Multi-View Retrieval for Knowledge-Based VQA
Hao Wang, Xujia Li, Lei Chen
Abstract
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieval from large-scale knowledge bases, yet this stage is often treated simplistically. Existing methods typically adopt single-view indexing or naive multi-view fusion, leading to systematic coverage gaps. In this work, we demonstrate that different views exhibit strong complementarity in retrieval. Motivated by this observation, we propose SCAR, a Structure-aware Cross-View Retrieval framework that exploits cross-view structural complementarity at inference time without additional training. SCAR enhances retrieval via structure-aware similarity propagation within each view and explicit cross-view redundancy regulation. Experiments on multiple KB-VQA benchmarks demonstrate that SCAR substantially improves retrieval recall, approaches retrieval coverage upper bounds, and consistently boosts end-to-end KB-VQA performance with negligible inference overhead 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesHexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal et al.ICCV 2023 · 123 citations
- Encyclopedic VQA: Visual questions about detailed properties of fine-grained categoriesThomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel et al.ICCV 2023 · 111 citations
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun et al.EMNLP 2023 · 37 citations
Related papers
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu et al.ICML 2026
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu et al.CVPR 2022 · 115 citations
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu et al.NeurIPS 2022 · 119 citations
- A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question AnsweringAlireza Salemi, Juan Altmayer Pizzorno, Hamed ZamaniSIGIR 2023 · 25 citations
