LLM Beliefs Are in Their Heads
Alessandro Corona Mendozza, Anders Søgaard
Abstract
We investigate belief-like representations in decoder-only autoregressive LLMs using linear controlled probes on residual stream activations and single attention heads. Following Herrmann and Levinstein's (2025) criteria (Accuracy, Use, Coherence, and Uniformity), we find that large models exhibit strong truthsensitivity (Accuracy), and steering activations along probe directions reliably changes downstream behavior (Use). Coherence, measured via calibrated probes and cross-dataset probing, is moderate across models, while training on diverse data yields domain-consistent truth directions (Uniformity). The results are particularly encouraging at the head level and align with some standard philosophical accounts of belief, e.g., minimal functionalism, supporting the view that LLMs can maintain propositional attitudes under such theoretical frameworks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c83bbcdf-f390-4eb5-9c97-acd290ccab72Builds on6
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Truth is Universal: Robust Detection of Lies in LLMsLennart Bürger, Fred A. Hamprecht, Boaz NadlerNeurIPS 2024 · 93 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
Related papers
- A Behavioural and Representational Evaluation of Goal-Directedness in Language Model AgentsRaghu Arghal, Fade Chen, Niall Dalton, Evgenii Kortukov et al.ICML 2026
- Monitoring Latent World States in Language Models with Propositional ProbesJiahai Feng, Stuart Russell, Jacob SteinhardtICLR 2025
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMsShivam Adarsh, Maria Maistro, Christina LiomaACL 2026
- Training-free Truthfulness Detection via Sparse MLP Value VectorsRunheng Liu, Heyan Huang, Xingchen Xiao, Yanghao Zhou et al.KDD 2026
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
