Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
Zongmeng Zhang, Yufeng Shi, Jinhua Zhu, Wengang Zhou, Xiang Qi, Peng Zhang, Houqiang Li
Abstract
Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers from hallucinations, one primary cause of which is the conflict between contextual and parametric knowledge. We deem that retrievalaugmented language models have the inherent capabilities of supplying response according to both contextual and parametric knowledge. Inspired by aligning language models with human preference, we take the first step towards aligning retrievalaugmented language models to a status where it responds relying merely on the external evidence and disregards the interference of parametric knowledge. Specifically, we propose a reinforcement learning based algorithm TRUSTWORTHY-ALIGNMENT, theoretically and experimentally demonstrating large language models' capability of reaching a trustworthy status without explicit supervision on how to respond. Our work highlights the potential of large language models on exploring its intrinsic abilities by its own and expands the application scenarios of alignment from fulfilling human preference to creating trustworthy agents. Our code is available at https://github.com/zmzhang2000/ trustworthy-alignment .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c34c73c-8b08-47cd-9771-7e528a26d298Cited by top-tier papers3
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive DomainsYash Saxena, Ankur Padia, Mandar Chaudhary, Kalpa Gunaratna et al.ICML 2026 · 7 citations
- OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language ModelsJunda Wu, Xintong Li, Ruoyu Wang, Yu Xia et al.ICLR 2025
- Robust Multimodal Large Language Models Against Modality ConflictZongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang LiICML 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveBowen Jin, Jinsung Yoon, Zhen Qin, Ziqi Wang et al.ICML 2025
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong et al.NeurIPS 2024 · 63 citations
- Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented GenerationGuanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang et al.WWW 2025 · 44 citations
- More RLHF, More Trust? On The Impact of Preference Alignment On TrustworthinessAaron Jiaxun Li, Satyapriya Krishna, Himabindu LakkarajuICLR 2025
- KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language ModelsRuizhe Zhang, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu et al.AAAI 2025 · 15 citations
