Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction Detection
Yuhao Chen, Yuanjie Lyu, Shuochen Liu, Chao Zhang, Junhui Lv, Tong Xu
Abstract
Detecting self-contradictions within documents is a challenging task for ensuring textual coherence and reliability. While large language models (LLMs) have advanced in many natural language understanding tasks, document-level self-contradiction detection (DSCD) remains insufficiently studied. Recent approaches leveraging Chain-of-Thought (CoT) prompting aim to enhance reasoning and interpretability; however, they only gain marginal improvement and often introduce inconsistencies across repeated responses. We observe that such inconsistency arises from incomplete reasoning chains that fail to include all relevant contradictory sentences consistently. To address this, we propose a two-stage method that combines supervised fine-tuning (SFT) and reinforcement learning (RL) to enhance DSCD performance. In the SFT phase, a teacher model helps the model learn reasoning patterns, while RL further refines its reasoning ability. Our method incorporates a task-specific reward function to expand the model's reasoning scope, boosting both accuracy and consistency. On the Con-traDoc benchmark, our approach significantly boosts Llama 3.1-8B-Instruct's accuracy from 38.5% to 51.1%, and consistency from 59.6% to 76.2%. 1 * Corresponding author. 1 Data and Code: https://github.com/isyuhaochen/RRC-DSCD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71e0fc79-ae37-4e44-9ac5-ddd79a1f4ebeBuilds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho et al.NeurIPS 2024 · 287 citations
- CDConv: A Benchmark for Contradiction Detection in Chinese ConversationsChujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng et al.EMNLP 2022 · 6 citations
Related papers
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-TuningWenqiao Zhu, Ji Liu, Rongjunchen Zhang, Haipang Wu et al.EMNLP 2025
- Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi et al.ACL 2025 · 13 citations
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-TuningQi (Cheems) Wang, Yanrui Yu, Ye Yuan, Rui Mao et al.NeurIPS 2025 · 103 citations
- Learning to Reason for Hallucination Span DetectionHsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna et al.ICLR 2026 · 8 citations
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
