Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Thomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araújo, Vittorio Ferrari
Abstract
We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It contains 221k unique question+answer pairs each matched with (up to) 5 images, resulting in a total of 1M VQA samples. Moreover, our dataset comes with a controlled knowledge base derived from Wikipedia, marking the evidence to support each answer. Empirically, we show that our dataset poses a hard challenge for large vision+language models as they perform poorly on our dataset: PaLI [14] is state-of-the-art on OK-VQA [37], yet it only achieves 13.0% accuracy on our dataset. Moreover, we experimentally show that progress on answering our encyclopedic questions can be achieved by augmenting large models with a mechanism that retrieves relevant information from the knowledge base. An oracle experiment with perfect retrieval achieves 87.0% accuracy on the single-hop portion of our dataset, and an automatic retrieval-augmented prototype yields 48.8%. We believe that our dataset enables future research on retrieval-augmented vision+language models. It is available at https://github.com/ google-research/google-research/tree/ master/encyclopedic_vqa .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers48
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang et al.ICLR 2026 · 79 citations
- Understanding Information Storage and Transfer in Multi-Modal Large Language ModelsSamyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi et al.NeurIPS 2024 · 57 citations
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun et al.EMNLP 2023 · 37 citations
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and FilteringYuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan et al.NeurIPS 2025 · 25 citations
- Towards General Continuous Memory for Vision-Language ModelsWenyi Wu, Zixuan Song, Kun Zhou, Yifei Shao et al.NeurIPS 2025 · 19 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringDuong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le PhuocICCV 2025 · 2 citations
- Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question AnsweringJialin Wu, Raymond J. MooneyEMNLP 2022 · 9 citations
- AraVQA: Building a New Arabic Factoid Visual Question Answering Dataset from WikipediaSultan Alrowili, Younes Samih, Abed Alhakim Freihat, Mathan Kumar EswaranACL 2026
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang et al.AAAI 2025 · 18 citations
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao et al.ACL 2026
