KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models
Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, Yulia Tsvetkov
摘要
Large language models (LLMs) demonstrate remarkable performance on knowledge-intensive tasks, suggesting that real-world knowledge is encoded in their model parameters. However, besides explorations on a few probing tasks in limited knowledge domains, it is not well understood how to evaluate LLMs' knowledge systematically and how well their knowledge abilities generalize, across a spectrum of knowledge domains and progressively complex task formats. To this end, we propose KGQuiz, a knowledge-intensive benchmark to comprehensively investigate the knowledge generalization abilities of LLMs. KGQuiz is a scalable framework constructed from triplet-based knowledge, which covers three knowledge domains and consists of five tasks with increasing complexity: true-or-false, multiple-choice QA, blank filling, factual editing, and open-ended knowledge generation. To gain a better understanding of LLMs' knowledge abilities and their generalization, we evaluate 10 open-source and black-box LLMs on the KGQuiz benchmark across the five knowledge-intensive tasks and knowledge domains. Extensive experiments demonstrate that LLMs achieve impressive performance in straightforward knowledge QA tasks, while settings and contexts requiring more complex reasoning or employing domain-specific facts still present significant challenges. We envision KGQuiz as a testbed to analyze such nuanced variations in performance across domains and task formats, and ultimately to understand, evaluate, and improve LLMs' knowledge abilities across a wide spectrum of knowledge domains and tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- KnowTuning: Knowledge-aware Fine-tuning for Large Language ModelsYougang Lyu, Lingyong Yan, Shuaiqiang Wang, Haibo Shi 等EMNLP 2024 · 被引用 3 次
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model AlignmentAndrey Sakhovskiy, Elena TutubalinaSIGIR 2025 · 被引用 2 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Deep Bidirectional Language-Knowledge Graph PretrainingMichihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang 等NeurIPS 2022 · 被引用 294 次
相关 Paper
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin 等EMNLP 2023 · 被引用 21 次
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue 等CVPR 2026
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen 等ACL 2026 · 被引用 1 次
- Explore What LLM Does Not Know in Complex Question AnsweringXin Lin, Zhenya Huang, Zhiqiang Zhang, Jun Zhou 等AAAI 2025 · 被引用 8 次
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 被引用 2 次
