HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation Editing
Tongxu Lin, Junping Du, Zhe Xue, Meiyu Liang, Runqing Tang
摘要
Large language models (LLMs) often generate hallucinations, undermining the reliability of their outputs. While prior work improves truthfulness by contrastive decoding or representation editing, these approaches overlook the hierarchical relationship between truthful and untruthful content, thereby constraining the LLM's knowledge potential. In this paper, we propose HyperEdit, a novel inference-time intervention method that encodes truthful and untruthful content as an entailment hierarchy and performs representation editing in the hyperbolic space to activate the truthfulness of LLMs. Specifically, we employ an auto-encoder to project the representations of LLMs into a hyperbolic space where truthful samples lie closer to the origin and untruthful ones farther away. This results in a truthfulness-aware hyperbolic space, where proximity to the origin indicates higher truthfulness, defining a natural editing direction. During inference, we edit the LLM's internal representations along this hyperbolic direction to elicit more truthful outputs. Experimental results on TruthfulQA and three additional benchmarks demonstrate that HyperEdit consistently enhances truthfulness across various advanced LLMs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen 等NeurIPS 2024 · 被引用 66 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- On the Universal Truthfulness Hyperplane Inside LLMsJunteng Liu, Shiqi Chen, Yu Cheng, Junxian HeEMNLP 2024
- Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMsGiovanni Servedio, Alessandro De Bellis, Dario Di Palma, Vito Walter Anelli 等ACL 2025 · 被引用 9 次
