Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
Tomer Ashuach, Shai Gretz, Yoav Katz, Yonatan Belinkov, Liat Ein-Dor
Abstract
Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether large language models possess similar privileged knowledge about answer correctness, information unavailable through external observation. We train correctness classifiers on question representations from both a model's own hidden states and external models, testing whether self-representations provide a performance advantage. On standard evaluation, we find no advantage: self-probes perform comparably to peer-model probes. We hypothesize this is due to high inter-model agreement of answer correctness. To isolate genuine privileged knowledge, we evaluate on disagreement subsets, where models produce conflicting predictions. Here, we discover domain-specific privileged knowledge: self-representations consistently outperform peer representations in factual knowledge tasks, but show no advantage in math reasoning. We further localize this domain asymmetry across model layers, finding that the factual advantage emerges progressively from early-to-mid layers onward, consistent with model-specific memory retrieval, while math reasoning shows no consistent advantage at any depth.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar et al.NeurIPS 2025 · 49 citations
- Looking Inward: Language Models Can Learn About Themselves by IntrospectionFelix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight et al.ICLR 2025 · 5 citations
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsJavier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel NandaICLR 2025 · 1 citation
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsHadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart et al.ICLR 2025
Related papers
- Estimating Knowledge in Large Language Models Without Generating a Single TokenDaniela Gottesman, Mor GevaEMNLP 2024 · 2 citations
- Calibrating Reasoning in Language Models with Internal ConsistencyZhihui Xie, Jizhou Guo, Tong Yu, Shuai LiNeurIPS 2024 · 37 citations
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 31 citations
- Teach Small Models to Reason by Curriculum DistillationWangyi Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han et al.EMNLP 2025
- Co-occurrence is not Factual Association in Language ModelsXiao Zhang, Miao Li, Ji WuNeurIPS 2024 · 15 citations
