Statistical Knowledge Assessment for Large Language Models
Qingxiu Dong, Jingjing Xu, Lingpeng Kong, Zhifang Sui, Lei Li
Abstract
Given varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we study the problem of quantifying knowledge contained in an LLM regarding a given set of facts. We propose KaRR, a statistical approach to assess factual knowledge for LLMs. The main idea is to estimate the ratio of LLM generating text corresponding to the answer entity given diverse prompts of the subject and the querying relation, versus it generating by random chances. Our assessment suite contains a comprehensive set of 994,123 entities and 600 relations, with 1,395,905 text aliases. We use our method to evaluate 20 LLMs of various sizes, including LLaMA, Alpaca, OPT, etc. Experiments show that our results have a strong correlation (0.43 Kendall's ) with the results of human assessment on LLMs. Our results reveal that the knowledge in LLMs with the same backbone architecture adheres to the scaling law, while tuning on instruction-following data sometimes compromises the model's capability to generate factually correct text reliably.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5e71fdd-4f1a-4171-80ca-6c0753333d1dCited by top-tier papers4
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Knowledge Conflicts for LLMs: A SurveyRongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang et al.EMNLP 2024 · 38 citations
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model EvaluationXunjian Yin, Xu Zhang, Jie Ruan, Xiaojun WanACL 2024
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeJiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng et al.ACL 2023 · 23 citations
- Editing Factual Knowledge in Language ModelsNicola De Cao, Wilker Aziz, Ivan TitovEMNLP 2021 · 20 citations
Related papers
- Factual Confidence of LLMs: on Reliability and Robustness of Current EstimatorsMatéo Mahaut, Laura Aina, Paula Czarnowska, Momchil Hardalov et al.ACL 2024 · 7 citations
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- The Effect of Scaling, Retrieval Augmentation and Form on the Factual Consistency of Language ModelsLovisa Hagström, Denitsa Saynova, Tobias Norlund, Moa Johansson et al.EMNLP 2023 · 6 citations
- Estimating Knowledge in Large Language Models Without Generating a Single TokenDaniela Gottesman, Mor GevaEMNLP 2024 · 2 citations
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language ModelsMert Yüksekgönül, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar et al.ICLR 2024 · 73 citations
