Beyond Single-View Indexing: Structure-Aware Multi-View Retrieval for Knowledge-Based VQA
Hao Wang, Xujia Li, Lei Chen
摘要
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieval from large-scale knowledge bases, yet this stage is often treated simplistically. Existing methods typically adopt single-view indexing or naive multi-view fusion, leading to systematic coverage gaps. In this work, we demonstrate that different views exhibit strong complementarity in retrieval. Motivated by this observation, we propose SCAR, a Structure-aware Cross-View Retrieval framework that exploits cross-view structural complementarity at inference time without additional training. SCAR enhances retrieval via structure-aware similarity propagation within each view and explicit cross-view redundancy regulation. Experiments on multiple KB-VQA benchmarks demonstrate that SCAR substantially improves retrieval recall, approaches retrieval coverage upper bounds, and consistently boosts end-to-end KB-VQA performance with negligible inference overhead 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesHexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal 等ICCV 2023 · 被引用 123 次
- Encyclopedic VQA: Visual questions about detailed properties of fine-grained categoriesThomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel 等ICCV 2023 · 被引用 111 次
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun 等EMNLP 2023 · 被引用 37 次
相关 Paper
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu 等ICML 2026
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu 等CVPR 2022 · 被引用 115 次
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan 等SIGIR 2026 · 被引用 2 次
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu 等NeurIPS 2022 · 被引用 119 次
- A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question AnsweringAlireza Salemi, Juan Altmayer Pizzorno, Hamed ZamaniSIGIR 2023 · 被引用 25 次
