HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation Editing
Tongxu Lin, Junping Du, Zhe Xue, Meiyu Liang, Runqing Tang
Abstract
Large language models (LLMs) often generate hallucinations, undermining the reliability of their outputs. While prior work improves truthfulness by contrastive decoding or representation editing, these approaches overlook the hierarchical relationship between truthful and untruthful content, thereby constraining the LLM's knowledge potential. In this paper, we propose HyperEdit, a novel inference-time intervention method that encodes truthful and untruthful content as an entailment hierarchy and performs representation editing in the hyperbolic space to activate the truthfulness of LLMs. Specifically, we employ an auto-encoder to project the representations of LLMs into a hyperbolic space where truthful samples lie closer to the origin and untruthful ones farther away. This results in a truthfulness-aware hyperbolic space, where proximity to the origin indicates higher truthfulness, defining a natural editing direction. During inference, we edit the LLM's internal representations along this hyperbolic direction to elicit more truthful outputs. Experimental results on TruthfulQA and three additional benchmarks demonstrate that HyperEdit consistently enhances truthfulness across various advanced LLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 20c4da86-af0a-4590-9b1b-a3fb184432edRelated papers
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen et al.NeurIPS 2024 · 66 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- On the Universal Truthfulness Hyperplane Inside LLMsJunteng Liu, Shiqi Chen, Yu Cheng, Junxian HeEMNLP 2024
- Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMsGiovanni Servedio, Alessandro De Bellis, Dario Di Palma, Vito Walter Anelli et al.ACL 2025 · 9 citations
