Lune

ICLR2026顶会

Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering

Mahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi Yu

出版方
2026年份

摘要

Recent advances in large language models (LLMs) have improved the accuracy of visual question answering (VQA) systems. However, directly applying LLMs to VQA still presents several challenges: (a) suboptimal performance when handling questions from specialized domains, (b) higher computational costs and slower inference speed due to large model sizes, and (c) the absence of a systematic approach to precisely quantify the uncertainty of LLM responses, raising concerns about their reliability in high-stakes tasks. To address these issues, we propose an UNcertainty-aware LLM-Integrated VQA model (Uni-VQA\texttt{Uni-VQA}). This model facilitates knowledge exchange between the LLM and a calibrated task-specific model (TS-VQA), guided by reliable confidence scores, resulting in improved VQA accuracy, reliability and inference speed. Our framework strategically leverages these confidence scores to manage the interaction between the LLM and TS-VQA\texttt{TS-VQA}: the specialized questions are answered by the TS-VQA\texttt{TS-VQA} model, while general knowledge questions are handled by the LLM. For questions requiring both specialized and general knowledge, the TS-VQA\texttt{TS-VQA} provides candidate answers, which the LLM then combines with its internal knowledge to generate a more accurate response. Extensive experiments on VQA datasets demonstrate the theoretically justified advantages of Uni-VQA\texttt{Uni-VQA} over using the LLM or TS-VQA\texttt{TS-VQA} alone.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext b6055f03-41be-4a30-a860-216d26fe4e7b

它引用的顶会 Paper28

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖