Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
Tomer Ashuach, Shai Gretz, Yoav Katz, Yonatan Belinkov, Liat Ein-Dor
摘要
Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether large language models possess similar privileged knowledge about answer correctness, information unavailable through external observation. We train correctness classifiers on question representations from both a model's own hidden states and external models, testing whether self-representations provide a performance advantage. On standard evaluation, we find no advantage: self-probes perform comparably to peer-model probes. We hypothesize this is due to high inter-model agreement of answer correctness. To isolate genuine privileged knowledge, we evaluate on disagreement subsets, where models produce conflicting predictions. Here, we discover domain-specific privileged knowledge: self-representations consistently outperform peer representations in factual knowledge tasks, but show no advantage in math reasoning. We further localize this domain asymmetry across model layers, finding that the factual advantage emerges progressively from early-to-mid layers onward, consistent with model-specific memory retrieval, while math reasoning shows no consistent advantage at any depth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar 等NeurIPS 2025 · 被引用 49 次
- Looking Inward: Language Models Can Learn About Themselves by IntrospectionFelix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight 等ICLR 2025 · 被引用 5 次
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsJavier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel NandaICLR 2025 · 被引用 1 次
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsHadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart 等ICLR 2025
相关 Paper
- Estimating Knowledge in Large Language Models Without Generating a Single TokenDaniela Gottesman, Mor GevaEMNLP 2024 · 被引用 2 次
- Calibrating Reasoning in Language Models with Internal ConsistencyZhihui Xie, Jizhou Guo, Tong Yu, Shuai LiNeurIPS 2024 · 被引用 37 次
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 被引用 31 次
- Teach Small Models to Reason by Curriculum DistillationWangyi Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han 等EMNLP 2025
- Co-occurrence is not Factual Association in Language ModelsXiao Zhang, Miao Li, Ji WuNeurIPS 2024 · 被引用 15 次
