Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
Shuzheng Si, Haozhe Zhao, Cheng Gao, Yuzhuo Bai, Zhitong Wang, Bofei Gao, Kangyang Luo, Wenhao Li, Yufei Huang, Gang Chen, Fanchao Qi, Minjia Zhang
Abstract
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable informationseeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58f9a7d8-2d81-454e-994e-bda7692b483eCited by top-tier papers5
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningLakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems et al.ICLR 2026 · 466 citations
- Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data FilteringShuzheng Si, Haozhe Zhao, Gang Chen, Cheng Gao et al.ACL 2025 · 11 citations
- Copy-Paste to Mitigate Large Language Model HallucinationsYongchao Long, Yingying Zhang, Xianbin Wen, Xian Wu et al.ICLR 2026 · 2 citations
- Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image GuidanceHaozhe Zhao, Shuzheng Si, Liang Chen, Yichi Zhang et al.EMNLP 2025 · 1 citation
- The Digital Dunning-Kruger Effect: Decoupling Hallucinations via Geometric Hidden-state Observation for Semantic TruthfulnessYueheng Mao, Min Yu, Gengwang Li, Jianguo Jiang et al.ACL 2026
Builds on16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationFangyuan Xu, Weijia Shi, Eunsol ChoiICLR 2024 · 260 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text GenerationSong Duong, Florian Le Bronnec, Alexandre Allauzen, Vincent Guigue et al.ICLR 2025
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang et al.ICML 2026
- Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced OptimizationLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan et al.ACL 2025
- Advancing Large Language Model Attribution through Self-ImprovingLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao et al.EMNLP 2024 · 2 citations
- Grounding by Trying: LLMs with Reinforcement Learning-Enhanced RetrievalSheryl Hsu, Omar Khattab, Chelsea Finn, Archit SharmaICLR 2025
