Explore What LLM Does Not Know in Complex Question Answering
Xin Lin, Zhenya Huang, Zhiqiang Zhang, Jun Zhou, Enhong Chen
摘要
Complex question answering (QA) is a challenging task in artificial intelligence research which requires reasoning based on related knowledge. The retrieval-augmented generation (RAG) based on large language models (LLMs) have become one promising solution in QA. To facilitate RAG more effectively, the LLM needs to precisely evaluate knowledge required in QA. That is, first, the LLM needs to examine its knowledge boundary (what the LLM does not know) to retrieve external knowledge as supplement. Second, the LLM needs to evaluate the utility of the retrieved knowledge (whether it helps in reasoning) for robust RAG. To this end, in this paper, we propose a novel Question Answering with Knowledge Evaluation (KEQA) framework to promote the effectiveness and efficiency of RAG in QA. First, inspired by quizzes in classroom, we propose a quiz-based method to precisely examine the knowledge state of the uninterpretable LLM for QA. We ask indicative quizzes on each required knowledge, and inspect whether the LLM can consistently answer the quiz to examine its knowledge boundary. Second, we retrieve the unknown knowledge from external source, and evaluate its utility to pick the helpful ones for reasoning. We design a reasoning-based metric to evaluate utility, and construct a demonstration set in training data for reference to guide knowledge picking in inference. We conduct extensive experiments on four widely-used QA datasets, and the results demonstrate the effectiveness of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale TablesShenghao Ye, Yu Guo, Dong Jin, Yuxiang Wang 等ACL 2026 · 被引用 8 次
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical ReasoningXilin Dang, Kexin Chen, Xiaorui Su, Ayush Noori 等ICLR 2026 · 被引用 6 次
- Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order FittingHaizhou Du, Jiujiu Li, Dongyang Li, Luobin Huang 等AAAI 2026
- Enhancing Pre-training Data Detection in LLMs Through Discriminative and Symmetric Prefix SelectionKai Sun, Yuxin Lin, Bo Dong, Jingyao Zhang 等AAAI 2026
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented GenerationZijun Yao, Weijian Qi, Liangming Pan, Shulin Cao 等ACL 2025
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
- Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back HomeViktor Moskvoretskii, Maria Marina, Mikhail Salnikov, Nikolay Ivanov 等ACL 2025 · 被引用 22 次
- Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based InferenceZhuo Chen, Xinyu Wang, Yong Jiang, Zhen Zhang 等EMNLP 2025 · 被引用 8 次
- REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question AnsweringYuhao Wang, Ruiyang Ren, Junyi Li, Xin Zhao 等EMNLP 2024 · 被引用 12 次
