UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
Boyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, Kam-Fai Wong
摘要
Despite demonstrating impressive capabilities, Large Language Models (LLMs) still often struggle to accurately express the factual knowledge they possess, especially in cases where the LLMs' knowledge boundaries are ambiguous. To improve LLMs' factual expressions, we propose the UAlign framework, which leverages Uncertainty estimations to represent knowledge boundaries, and then explicitly incorporates these representations as input features into prompts for LLMs to Align with factual knowledge. First, we prepare the dataset on knowledge question-answering (QA) samples by calculating two uncertainty estimations, including confidence score and semantic entropy, to represent the knowledge boundaries for LLMs. Subsequently, using the prepared dataset, we train a reward model that incorporates uncertainty estimations and then employ the Proximal Policy Optimization (PPO) algorithm for factuality alignment on LLMs. Experimental results indicate that, by integrating uncertainty representations in LLM alignment, the proposed UAlign can significantly enhance the LLMs' capacities to confidently answer known questions and refuse unknown questions on both in-domain and out-of-domain tasks, showing reliability improvements and good generalizability over various prompt- and training-based baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMsJunxiao Yang, Jinzhe Tu, Haoran Liu, Xiaoce Wang 等ICLR 2026 · 被引用 9 次
- Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM HallucinationsWen Luo, Guangyue Peng, Wei Li, Shaohang Wei 等ACL 2026 · 被引用 3 次
- From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty CalibrationHao Li, Tao He, Jiafeng Liang, Zheng Chu 等AAAI 2026
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
相关 Paper
- TUR-DPO: Topology- and Uncertainty-Aware Direct Preference OptimizationAbdulhady abas, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad OussalahICML 2026
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong 等NeurIPS 2024 · 被引用 63 次
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang 等ICML 2026
- Trust Within? Seek Beyond? Knowledge Boundary Aware Policy Optimization for Agentic SearchTao Feng, Xinke Jiang, Xinyan Hu, Yonggang Zhang 等ACL 2026 · 被引用 1 次
- UCPO: Uncertainty-Aware Policy OptimizationXianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan 等ICML 2026
