On the Universal Truthfulness Hyperplane Inside LLMs
Junteng Liu, Shiqi Chen, Yu Cheng, Junxian He
Abstract
While large language models (LLMs) have demonstrated remarkable abilities across various fields, hallucination remains a significant challenge. Recent studies have explored hallucinations through the lens of internal representations, proposing mechanisms to decipher LLMs' adherence to facts. However, these approaches often fail to generalize to out-of-distribution data, leading to concerns about whether internal representation patterns reflect fundamental factual awareness, or only overfit spurious correlations on the specific datasets. In this work, we investigate whether a universal truthfulness hyperplane that distinguishes the model's factually correct and incorrect outputs exists within the model. To this end, we scale up the number of training datasets and conduct an extensive evaluation -we train the truthfulness hyperplane on a diverse collection of over 40 datasets and examine its cross-task, cross-domain, and in-domain generalization. Our results indicate that increasing the diversity of the training datasets significantly enhances the performance in all scenarios, while the volume of data samples plays a less critical role. This finding supports the optimistic hypothesis that a universal truthfulness hyperplane may indeed exist within the model, offering promising directions for future research. Code is publicly available at https://github.com/hkust-nlp/ Universal_Truthfulness_Hyperplane . Tend to overfit TruthfulQA OOD In-distribution …… Reasoning QA Topic Classification Coreference Reading Comprehe nsion Others Common Sense QA Diverse Data Universal Truthfulness Hyperplane OOD Be ats music is owned by Pear Inc .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- Language Models Can Predict Their Own BehaviorDhananjay Ashok, Jonathan MayNeurIPS 2025 · 10 citations
- Prompt-Guided Internal States for Hallucination Detection of Large Language ModelsFujie Zhang, Peiqi Yu, Biao Yi, Baolei Zhang et al.ACL 2025 · 8 citations
- FLaG: Fine-Grained Latent Grouping for Hallucination DetectionWentao Ye, Liyao Li, Zhiqing Xiao, Muzhi Zhu et al.KDD 2026 · 1 citation
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetySeongmin Lee, Aeree Cho, Grace C. Kim, Shengyun Peng et al.EMNLP 2025 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMsGiovanni Servedio, Alessandro De Bellis, Dario Di Palma, Vito Walter Anelli et al.ACL 2025 · 9 citations
- HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation EditingTongxu Lin, Junping Du, Zhe Xue, Meiyu Liang et al.KDD 2026
- TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination via Latent Truthful-Guided Pre-InterventionJinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer et al.ICCV 2025
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsHadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart et al.ICLR 2025
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang et al.ICML 2026
