Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty
Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li
摘要
The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised fine-tuning (SFT), a common technique for model specialization. Existing recovery methods rely on data-intensive global parameter adjustments, implicitly assuming that SFT deeply corrupts the models' ability to recognize their knowledge boundaries. However, we observe that fine‑tuned LLMs still preserve this ability; what is damaged is their capacity to faithfully express that awareness. Building on this, we propose Honesty-Critical Neurons Restoration (HCNR) to surgically repair this suppressed capacity. HCNR identifies and restores key expression-governing neurons to their pre-trained state while harmonizing them with task-oriented neurons via Hessian-guided compensation. Experiments on four QA tasks and five LLM families demonstrate that HCNR effectively recovers 33.25% of the compromised honesty while achieving at least 2.23x speedup with over 10x less data compared to baseline methods, offering a practical solution for trustworthy LLM deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig 等NeurIPS 2024 · 被引用 82 次
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 被引用 71 次
- Confidence Regulation Neurons in Language ModelsAlessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov 等NeurIPS 2024 · 被引用 68 次
相关 Paper
- Alleviating the Fear of Losing Alignment in LLM Fine-tuningKang Yang, Guanhong Tao, Xun Chen, Jun XuS&P 2025
- HonestLLM: Toward an Honest and Helpful Large Language ModelChujie Gao, Siyuan Wu, Yue Huang, Dongping Chen 等NeurIPS 2024 · 被引用 30 次
- NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-TuningXin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo 等AAAI 2025 · 被引用 38 次
- Unlearners Can Lie: Evaluating and Improving Honesty in LLM UnlearningRenjie Gu, Jiazhen Du, Yihua Zhang, Sijia LiuACL 2026
- From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint TuningWei Chen, Zhen Huang, Liang Xie, Binbin Lin 等ICML 2024 · 被引用 55 次
