Combo of Thinking and Observing for Outside-Knowledge VQA
Qingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji, Weiping Wang
摘要
Outside-knowledge visual question answering is a challenging task that requires both the acquisition and the use of open-ended real-world knowledge. Some existing solutions draw external knowledge into the cross-modality space which overlooks the much vaster textual knowledge in natural-language space, while others transform the image into a text that further fuses with the textual knowledge into the natural-language space and completely abandons the use of visual features. In this paper, we are inspired to constrain the cross-modality space into the same space of natural-language space which makes the visual features preserved directly, and the model still benefits from the vast knowledge in natural-language space. To this end, we propose a novel framework consisting of a multimodal encoder, a textual encoder and an answer decoder. Such structure allows us to introduce more types of knowledge including explicit and implicit multimodal and textual knowledge. Extensive experiments validate the superiority of the proposed method which outperforms the state-ofthe-art by 6.17% accuracy. We also conduct comprehensive ablations of each component, and systematically study the roles of varying types of knowledge. Codes and knowledge data can be found at https://github.com/ PhoebusSi/Thinking-while-Observing . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalWei Yang, Jingjing Fu, Rui Wang, Jinyu Wang 等ACL 2025 · 被引用 11 次
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 被引用 5 次
- Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringQingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu 等EMNLP 2023 · 被引用 2 次
- KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question AnsweringZhiyang Li, Ao Ke, Yukun Cao, Xike XieACL 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu 等AAAI 2022 · 被引用 517 次
- PaLI: A Jointly-Scaled Multilingual Language-Image ModelXi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni 等ICLR 2023 · 被引用 194 次
相关 Paper
- Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question AnsweringFeng Gao, Qing Ping, Govind Thattai, Aishwarya N. Reganti 等CVPR 2022 · 被引用 85 次
- Multi-Modal Answer Validation for Knowledge-Based VQAJialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh MottaghiAAAI 2022 · 被引用 183 次
- A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQAYangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu 等ACM MM 2022 · 被引用 41 次
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu 等CVPR 2022 · 被引用 115 次
