Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering
Mahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi Yu
Abstract
Recent advances in large language models (LLMs) have improved the accuracy of visual question answering (VQA) systems. However, directly applying LLMs to VQA still presents several challenges: (a) suboptimal performance when handling questions from specialized domains, (b) higher computational costs and slower inference speed due to large model sizes, and (c) the absence of a systematic approach to precisely quantify the uncertainty of LLM responses, raising concerns about their reliability in high-stakes tasks. To address these issues, we propose an UNcertainty-aware LLM-Integrated VQA model (). This model facilitates knowledge exchange between the LLM and a calibrated task-specific model (TS-VQA), guided by reliable confidence scores, resulting in improved VQA accuracy, reliability and inference speed. Our framework strategically leverages these confidence scores to manage the interaction between the LLM and : the specialized questions are answered by the model, while general knowledge questions are handled by the LLM. For questions requiring both specialized and general knowledge, the provides candidate answers, which the LLM then combines with its internal knowledge to generate a more accurate response. Extensive experiments on VQA datasets demonstrate the theoretically justified advantages of over using the LLM or alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6055f03-41be-4a30-a860-216d26fe4e7bBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question AnsweringZhongfan Sun, Kan Guo, Yongli Hu, Daxin Tian et al.ACM MM 2025
- Ranked from Within: Ranking Large Multimodal Models Without LabelsWeijie Tu, Weijian Deng, Dylan Campbell, Yu Yao et al.ICML 2025
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 5 citations
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models ReasoningWenyi Xiao, Xinchi Xu, Leilei GanACL 2026 · 2 citations
- CER: Confidence Enhanced Reasoning in LLMsAli Razghandi, Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani BaghshahACL 2025 · 11 citations
