Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, Soujanya Poria
摘要
LLMs are an integral component of retrieval-augmented generation (RAG) systems. While many studies focus on evaluating the overall quality of end-to-end RAG systems, there is a gap in understanding the appropriateness of LLMs for the RAG task. To address this, we introduce TRUST-SCORE, a holistic metric that evaluates the trustworthiness of LLMs within the RAG framework. Our results show that various prompting methods, such as in-context learning, fail to effectively adapt LLMs to the RAG task as measured by TRUST-SCORE. Consequently, we propose TRUST-ALIGN, a method to align LLMs for improved TRUST-SCORE performance. 26 out of 27 models aligned using TRUST-ALIGN substantially outperform competitive baselines on ASQA, QAMPARI, and ELI5. Specifically, in LLaMA-3-8b, TRUST-ALIGN outperforms FRONT on ASQA (↑12.56), QAM-PARI (↑36.04), and ELI5 (↑17.69). TRUST-ALIGN also significantly enhances models' ability to correctly refuse and provide quality citations. We also demonstrate the effectiveness of TRUST-ALIGN across different open-weight models, including the LLaMA series (1b to 8b), Qwen-2.5 series (0.5b to 7b), and Phi3.5 (3.8b). We release our code at https://github.com/declare-lab/ trust-align.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang 等NeurIPS 2025 · 被引用 85 次
- Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented GenerationRuizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao 等ICLR 2026 · 被引用 11 次
- Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAGXin Sun, Jianan Xie, Zhongqi Chen, Qiang Liu 等ACL 2025 · 被引用 11 次
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 被引用 10 次
- OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAGFengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang 等WWW 2026 · 被引用 8 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- CrAM: Credibility-Aware Attention Modification in LLMs for Combating Misinformation in RAGBoyi Deng, Wenjie Wang, Fengbin Zhu, Qifan Wang 等AAAI 2025 · 被引用 21 次
- GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal SynthesisYi Jiang, Sendong Zhao, Jianbo Li, Haochun Wang 等ACL 2025
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
- Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented GenerationShicheng Xu, Liang Pang, Mo Yu, Fandong Meng 等ACL 2024
- CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented ValidationYee Man Choi, Xuehang Guo, Yi R. Fung, Qingyun WangACL 2026 · 被引用 7 次
