BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
MinJun Kim, Seungwoo Song, Youhan Lee, Haneol Jang, Kyungtae Lim
Abstract
The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circumstances, the demand for multilingual evaluation of visual question answering (VQA) tasks, a representative task of multimodal systems, has increased. Accordingly, we propose a bilingual outside-knowledge VQA (BOK-VQA) dataset in this study that can be extended to multilingualism. The proposed data include 17K images, 17K question-answer pairs for both Korean and English and 280K instances of knowledge information related to question-answer content. We also present a framework that can effectively inject knowledge information into a VQA system by pretraining the knowledge information of BOK-VQA data in the form of graph embeddings. Finally, through in-depth analysis, we demonstrated the actual effect of the knowledge information contained in the constructed training data on VQA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6b74875-c5f6-4bcf-bb57-5a95d601cc65Builds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank et al.ACL 2022 · 147 citations
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QASeongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo et al.AAAI 2021 · 64 citations
- CARETS: A Consistency And Robustness Evaluative Test Suite for VQACarlos E. Jimenez, Olga Russakovsky, Karthik NarasimhanACL 2022
Related papers
- Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringZhen Yang, Zhuo Tao, Qi Chen, Liang Li et al.CVPR 2025
- Location-Aware Visual Question Generation with Lightweight ModelsNicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang et al.EMNLP 2023 · 3 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- Combo of Thinking and Observing for Outside-Knowledge VQAQingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji et al.ACL 2023 · 7 citations
- KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQAKenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta et al.CVPR 2021
