QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models
Wei Wang, Zhaowei Li, Qi Xu, Yiqing Cai, Hang Song, Qi Qi, Ran Zhou, Zhida Huang, Tao Wang, Li Xiao
摘要
The deployment of large language models (LLMs) faces considerable challenges concerning resource constraints and inference efficiency. Recent research has increasingly focused on smaller, task-specific models enhanced by distilling knowledge from LLMs. However, prior studies have often overlooked the diversity and quality of knowledge, especially the untapped potential of negative knowledge. Constructing effective negative knowledge remains severely understudied. In this paper, we introduce a novel framework called quality-guided contrastive rationale distillation aimed at enhancing reasoning capabilities through contrastive knowledge learning. For positive knowledge, we enrich its diversity through temperature sampling and employ selfconsistency for further denoising and refinement. For negative knowledge, we propose an innovative self-adversarial approach that generates low-quality rationales by sampling previous iterations of smaller language models, embracing the idea that one can learn from one's own weaknesses. A contrastive loss is developed to distill both positive and negative knowledge into smaller language models, where an online-updating discriminator is integrated to assess qualities of rationales and assign them appropriate weights, optimizing the training process. Through extensive experiments across multiple reasoning tasks, we demonstrate that our method consistently outperforms existing distillation techniques, yielding higher-quality rationales. Our codes will be released soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual ScenesChuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu 等ACL 2026 · 被引用 9 次
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language ModelsPu Jian, Junhong Wu, Wei Sun, Chen Wang 等EMNLP 2025 · 被引用 2 次
- MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT DistillationJin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang 等ACL 2026 · 被引用 1 次
- Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal ModelsWei Wang, Zhaowei Li, Qi Xu, Linfeng Li 等EMNLP 2025 · 被引用 1 次
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper6
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal 等ICML 2023 · 被引用 347 次
相关 Paper
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen 等EMNLP 2024 · 被引用 1 次
- Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative DataYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan 等AAAI 2024 · 被引用 31 次
- The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical ReasoningHaolong Qian, Xianliang Yang, Ma yinuo, Lirong Che 等ICML 2026
- MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language ModelsJustin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, Mohit BansalICML 2024 · 被引用 32 次
- Learning from Diverse Reasoning Paths with Routing and CollaborationZhenyu Lei, Zhen Tan, Song Wang, Yaochen Zhu 等EMNLP 2025
